A photograph of a cat can look reassuringly ordinary. The ears sit a little differently, the eyes narrow, the muzzle changes. To a trained observer, small details may contribute to a pain assessment. To a phone camera, they are pixels.
Researchers are trying to bridge that gap. The work is promising, but “AI can detect pain” combines several different claims. A system trained specifically on feline faces is not the same tool as a general chatbot invited to inspect a photograph.
The scale comes before the software
The Feline Grimace Scale was developed and validated for acute pain assessment. It considers five features: ear position, narrowing around the eyes, muzzle tension, whisker changes and head position.
That provides a structured framework rather than a vague judgement that a cat “looks sad”. It also defines the task researchers are asking a model to perform. Recognising a feline face and assessing signs associated with pain are separate achievements.
A single image still has limits. A score is not a diagnosis of the underlying cause, and a face is only part of an animal’s condition. The scale’s own resources are a better place to understand its intended use than an unlabelled graphic circulating online.
What a specialist model was trained to do
In a 2023 study, researchers annotated 3,447 cat-face images with 37 landmarks. Their system combined a neural network for locating facial features with models that used geometric information to predict pain scores.
The best reported classification accuracy was 95.5% under the study’s evaluation conditions. The authors presented the work as a basis for developing smartphone assessment technology.
The percentage is interesting, but it needs its surrounding sentence. It does not mean that any app will correctly assess any cat in any photograph 95.5% of the time. The data, labels, test design and population define what the number describes.
A chatbot faces a different test
A study published in 2025 compared four chatbots with an expert veterinarian using 50 cat-face images and repeated testing after two months. It examined agreement and the spread of individual errors, rather than accepting fluent explanations as evidence of competence.
The results raised concerns about underestimation and inconsistent agreement. Even an acceptable average difference could coexist with errors large enough to matter for an individual animal. These findings apply to the versions and conditions tested; they are not a permanent ranking of every future product.
This is why a convincing paragraph and a well-validated measurement should not be confused. A chatbot can describe the correct facial features while still assigning an unreliable score to the image in front of it.
Ask what happens when the model is wrong
A missed painful cat and an unnecessary alarm have different consequences. An evaluation should make both visible. It should also explain how the system performs on cats and conditions outside its development data, not merely give one attractive headline figure.
For caregivers, a reassuring automated answer should never delay veterinary advice when a cat’s behaviour, eating or movement has changed. Do not use a photo score to choose pain medication. The practical question is whether an animal needs professional assessment, not whether a screen can produce a number.
Our history of AI learning to recognise cats begins with finding a familiar visual category. Pain assessment asks considerably more of the technology. It has to be reliable precisely when the difference is subtle and the cost of missing it belongs to a living animal.




