2026 - 2026 · Animal Behaviour

Do Cat-egories Really Exist?

Researcher[s]: Guan, T.; ...;

About

This project investigated whether personality variation in domestic cats is better interpreted as a set of discrete behavioural profiles or as continuous multidimensional variation. Using established personality-factor scores from 4,316 cats in a publicly available Finnish owner-questionnaire dataset, the study evaluated Gaussian mixture representations across five core personality dimensions: fearfulness, activity/playfulness, aggression toward humans, sociability toward humans, and sociability toward cats. Rather than treating the output of a clustering algorithm as evidence of natural categories, the analysis separated density fit, classification quality, resampling reproducibility, cross-method agreement, demographic sensitivity, and sensitivity to the dimensions included in the personality space. The objective was to determine whether apparently stable feline personality profiles remain convincing when tested against plausible continuous alternatives and multiple analytical assumptions.

Research Focus

The central question was whether a statistically reproducible partition should be interpreted as evidence of distinct feline personality types. The analysis therefore distinguished three related but non-equivalent questions: whether mixture models improve the description of the observed personality-score distribution, whether the resulting assignments recur under resampling, and whether the same profile structure persists across alternative clustering methods, demographic adjustment, and changes in the behavioural dimensions included.

Data and Personality Space

The analysis used Version 2 of the publicly released Finnish feline behaviour and personality dataset, containing 4,316 domestic cats. Five established factor scores formed the primary personality space: fearfulness, activity/playfulness, aggression toward humans, sociability toward humans, and sociability toward cats. Excessive grooming and litterbox issues were reserved for a secondary seven-factor sensitivity analysis because they represent problematic behaviours rather than the same core personality dimensions. All released cats had complete scores for both the five- and seven-factor spaces, allowing the primary and expanded analyses to use identical individuals.

Study Design

All personality factors were standardized before modelling. Gaussian mixture models with one to eight components were fitted using both full- and diagonal-covariance structures. Bayesian information criterion (BIC) was used to identify the strongest density representation, while integrated completed likelihood (ICL) additionally penalized uncertain classification and was used as the reference profile representation. This distinction allowed density approximation and profile interpretability to be evaluated separately rather than assuming that the best-fitting density automatically defines meaningful personality categories.

Evidence for discreteness was then assessed against 200 continuous Gaussian-copula reference datasets constructed to preserve the empirical marginal distributions and approximate the dependence structure among personality traits without introducing latent classes. Bootstrap resampling of 500 datasets evaluated recovery of the preferred component count and assignment reproducibility, while 100 repeated split-half analyses tested whether independently fitted subsets of cats generated comparable partitions when projected back onto the same reference population. Adjusted Rand index (ARI), posterior assignment probabilities, entropy, centroid correspondence, and silhouette statistics were used to quantify different aspects of stability and separation.

Robustness was further assessed by comparing the Gaussian-mixture partition with k-means and Ward hierarchical clustering, residualizing personality factors for age, sex, and breed group, and repeating the analysis after adding excessive grooming and litterbox issues to produce a seven-factor behavioural space. These sensitivity analyses tested whether the apparent profile structure reflected a stable property of the cats or depended substantially on modelling geometry, demographic composition, or the definition of the behavioural feature space.

Analytical Framework

The analytical framework deliberately separated four forms of evidence:

  • Density structure: whether multiple Gaussian components described the observed personality-score distribution better than a single component.

  • Geometric separation: whether fitted components formed clearly separated regions of personality space rather than overlapping density approximations.

  • Internal reproducibility: whether component counts and individual assignments recurred under bootstrap and independent split-half resampling.

  • Analytical robustness: whether the same profiles persisted across alternative clustering algorithms, demographic adjustment, and expanded behavioural dimensions.

A profile solution was considered substantively convincing only if these lines of evidence converged, rather than because any single clustering criterion returned multiple components.

Key Findings

BIC and ICL gave markedly different answers. BIC selected an eight-component diagonal-covariance mixture at the upper boundary of the search range, indicating that increasingly complex mixtures continued to improve density approximation. In contrast, ICL selected a two-component full-covariance representation that provided a simpler and more confidently classifiable profile summary. The two components broadly contrasted cats showing lower fearfulness and aggression with higher activity and sociability against cats showing the opposite pattern.

The observed improvement in density fit exceeded that of the continuous Gaussian-copula reference datasets (Monte Carlo p = 0.00995). However, the BIC-selected density partition had a silhouette of only 0.046-lower than every continuous-reference value-showing that stronger mixture fit did not translate into unusually clear geometric separation. Thus, the data contained structure not fully captured by the continuous reference, but that structure did not provide straightforward evidence of discrete personality groups.

The two-component ICL representation was highly reproducible within the Gaussian-mixture framework. Bootstrap resampling recovered the two-component solution in 89.0% of samples, and fixed-model bootstrap assignments achieved a median ARI of 0.932. Independently fitted split halves recovered the same component count in 96.0% of half-sample fits, with a median partition ARI of 0.861. Mean maximum posterior assignment probability was 0.911, and 84.6% of cats had maximum posterior probability of at least 0.80.

That internal reproducibility did not extend equally across analytical choices. Agreement between the Gaussian-mixture partition and alternative clustering methods was substantially weaker, with ARI = 0.364 for k-means and 0.354 for Ward clustering. Demographic residualization retained the broad two-component contrast and produced highly correlated matched centroids (r = 0.992), yet individual assignment agreement fell to ARI = 0.593. When excessive grooming and litterbox issues were added, ICL selected four components, but bootstrap recovery fell to 45.8% and assignment stability weakened substantially.

Main Interpretation

The results support a reproducible mixture-based summary of feline personality variation, but not robust evidence for a natural taxonomy of discrete personality types. A stable algorithmic boundary can recur under resampling even when alternative, defensible analytical methods divide the same personality space differently. The distinction between repeatability and discreteness is therefore central: the two-component representation is useful as a descriptive shorthand, but its boundaries should not be treated as fixed behavioural identities.

For behavioural interpretation, retaining the underlying trait scores preserves information that categorical labels can conceal. Two cats assigned to the same profile may still differ substantially in fearfulness, activity, or sociability, while two cats on opposite sides of a fitted boundary may remain behaviourally similar. The study therefore favours dimensional descriptions when individual behavioural assessment is the goal.

Key Methods

Gaussian mixture modelling; model-based clustering; Bayesian information criterion (BIC); integrated completed likelihood (ICL); Gaussian-copula simulation; Monte Carlo reference testing; bootstrap resampling; split-half reproducibility analysis; adjusted Rand index; posterior classification probabilities; entropy analysis; silhouette analysis; principal component analysis; k-means clustering; Ward hierarchical clustering; demographic residualization; natural cubic splines; sensitivity analysis; Hungarian centroid matching.

Keywords

Domestic cats; feline personality; animal behaviour; personality profiles; behavioural phenotyping; Gaussian mixture models; model-based clustering; cluster stability; continuous traits; bootstrap reproducibility; Gaussian copula; behavioural classification; sensitivity analysis;

Project Status

Completed (2026)

← Back to Research