Can welfare metrics show what helps shelter dogs? A four-week study in context
Table of contents
A source-backed Methods Note about telemetry, enrichment and why one physiological number cannot stand in for a dog's whole welfare experience.
Research & Evidence | Methods Note | Updated September 4, 2026
The short answer
A study of eight intact male mixed-breed dogs in an Italian shelter recorded heart rate, body temperature and muscular activity every minute for four weeks, or 28 days. After a baseline week, the dogs experienced three one-week conditions: objects in the pen, a familiar human visitor and the presence of a female dog. The researchers used those signals to calculate three measures: predictability, sleep quality and the day-night cycle.
The social conditions showed the clearest changes in sleep quality and day-night cyclicity in this small group. The study also found large differences between individual dogs. Its predictability measure fell whenever the environment changed, but the authors explain that this only shows a change from the baseline pattern. It does not, by itself, say that welfare improved or worsened.
This brief asks: What did the study's metrics measure, and why should they not be treated as a universal welfare score for dogs?
What the study measured
Travain and colleagues worked with eight intact, healthy, mixed-breed male dogs at a private shelter in Bracciano, near Rome, Italy. The dogs were estimated to be younger than five years old. Six were medium-sized and two were large-sized. Each dog lived alone in a pen, and the dogs had been in the shelter for between three months and two years when the study began.
The researchers used an implanted telemetry transmitter. It recorded three signals: heart rate in beats per minute, body temperature in degrees Celsius and muscular activity in movement units. The system produced one mean reading per minute, around the clock. The study therefore collected 40,320 readings per dog during the 28 days of the experiment, or 322,560 readings across the eight dogs.
The implantation matters when reading the paper. The researchers first recorded a two-week post-operative period to monitor the dogs. They discarded those 14 days from the analysis. The four-week experiment began afterward. This was a research protocol that required surgery and a receiver system in the pens, not a household tracking routine.
The four-week comparison
All dogs began with one week of standard housing as a baseline. They then experienced three one-week conditions. The order varied between dogs, although the authors say that the small sample and external constraints prevented complete randomization.
| Study condition | What the dogs experienced | What the comparison can tell us |
|---|---|---|
| Baseline | The dog lived alone with its usual platform, food bowl and water bowl. | Each dog's own baseline pattern for comparison. |
| Objects | A cot, balls, a rubber chew toy, food, a scented rag and a natural beef bone were placed in the pen. | How this bundle of inanimate changes coincided with the signals. |
| Human presence | The same familiar female volunteer spent two hours with each dog once a day for that week's visits. | How a repeated human visit coincided with the signals. |
| Female dog | An unknown spayed female shared the pen after a compatibility check. | How this specific social arrangement coincided with the signals. |
These were not three interchangeable products or three universal treatments. They combined different kinds of social and environmental change, and each condition lasted one week. A result from this design cannot tell us which single object caused a change, whether a different person would have the same effect or how another dog would respond.
Three metrics, three questions
The paper calls its measures welfare metrics, but the measures do not mean the same thing.
| Metric | What it measured | What the study reported | Safe reading |
|---|---|---|---|
| Predictability | How similar each day's physiological pattern was to an individual average day built from the baseline week. | Predictability was lower during all three enrichment weeks. The differences from baseline were statistically significant at p=0.031. | The dog's measured pattern changed. A lower value is not automatically better or worse. |
| Sleep quality | Nighttime sleep and awakenings, including how often the dog woke and the intervals and length of those awakenings. | Human presence and the female-dog condition increased the measure. Only the female-dog condition was statistically different from baseline at p=0.030; the other comparison was not statistically significant at p=0.096. | A change in sleep-related signals is one part of the picture, not a complete welfare verdict. |
| Day-night cycle | The difference between daytime and nighttime heart rate and muscular activity. | Human presence and the female-dog condition were statistically higher than baseline at p=0.039 and p=0.046. The objects condition was not statistically different from baseline at p=0.159. | The signals showed a different day-night pattern in this sample. That does not define a universal ideal. |
The authors also report that the social conditions had stronger effects on sleep quality and cyclicity than the objects condition. They describe the female-dog condition as the strongest of the three for both measures in this study. That conclusion belongs to these eight dogs, this shelter, this sequence and these measurements.
Why predictability needs caution
Predictability is easy to misread. A measure can tell us that the dog's physiology moved away from its usual baseline pattern. It cannot tell us whether the new pattern is pleasant, unpleasant or simply a normal response to something different.
The study makes this point in its discussion. The predictability value fell when an enrichment condition changed the dogs' environment, but that metric had no built-in direction for positive or negative welfare. The change could show that the dogs noticed the new condition. It could not tell the researchers what the experience meant to each dog without other information.
That is why the paper's sleep and day-night results need to be read separately. They use different parts of the data and answer different questions. A single combined number would hide those distinctions.
What the results do and do not establish
The study supports the idea that continuous physiological data can reveal patterns that are difficult to see from occasional observations. It also shows why individual response matters. The dogs did not react as one uniform group, and the authors note that little was known about their previous lives or individual personalities.
The study does not establish that:
- a toy, visitor or companion dog will improve every dog's welfare;
- a lower or higher physiological value is a direct readout of happiness, distress or health;
- the three measures form a validated welfare score for household dogs;
- the female-dog condition is the best choice for every shelter or every dog; or
- a wearable or home device can reproduce this implanted research protocol.
The study's abstract uses a broad welfare conclusion, but its detailed results are more specific. Predictability changed under all three conditions, while the clearest sleep and day-night changes appeared under the social conditions. The authors also report substantial individual variation. Those details are part of the result, not footnotes to remove.
What the methods companion adds
A 2025 critical review by Cobb, Jiménez and Dreschel addresses the wider problem of physiological welfare measurement. It argues that researchers should not rely on isolated indicators, especially cortisol alone, to represent a dog's welfare. The review calls for multiple physiological indicators, attention to individual differences and clear reporting of characteristics such as age, body weight and sex.
That review does not validate the shelter study's metrics, a VEP score or a household intervention. It helps explain why the shelter paper is best read as a methods example. The useful question is not, "Which number is the dog?" It is, "What did this measure capture, what else was observed, and what remains uncertain?"
Five evidence-led insights
The study offers more than a list of measurements. It also shows how to read evidence when a number is precise but its meaning is limited. These are bounded inferences from the study and the methods review, not new results or a welfare score.
1. A changed signal is not a welfare verdict
Predictability changed during each enrichment week, but the paper does not give that change a positive or negative direction. That distinction is easy to lose when one number looks precise.
What this changes: Treat a physiological shift as evidence that a pattern changed. It is not a shortcut to "better," "worse," "happy" or "distressed."
2. Social context stood out, but the finding is not a universal ranking
Human presence and the female-dog condition showed the clearest changes in sleep quality and day-night cyclicity in this sample. Each condition was still a bundle of social and environmental changes, and the comparison involved eight dogs in one shelter.
What this changes: When considering an enrichment idea, ask what kind of change it introduces for this dog and setting. Do not turn the female-dog condition into a general recommendation.
3. The individual baseline is part of the result
The study built predictability around each dog's baseline and reported substantial variation between dogs. The methods review makes the same broader point: isolated physiological indicators can mislead without individual context.
What this changes: Compare a change with the dog's usual pattern and other observations. Do not treat a raw metric value as directly comparable between dogs or as a complete account of one dog's experience.
4. The protocol tested packages, not single causes
The objects condition combined several items, the human condition combined a person and a repeated schedule, and the social condition changed the dog's companion. The order was not completely randomized. Those choices shape what the comparison can say.
What this changes: A result can show that a bundled condition coincided with a changed pattern. It cannot identify one object, person or interaction as the cause.
5. More readings do not automatically mean a complete welfare picture
The telemetry system collected a detailed record, but detail in one physiological channel does not cover every part of welfare. The companion review supports using multiple indicators and attending to characteristics such as age, body weight, sex and context.
What this changes: Use numbers as one part of a wider observation record. Behavior, routine, environment and professional assessment still matter when a change is concerning.
A useful boundary for pet families
If you keep a behavior or activity log for your dog, treat each number as one observation in a larger pattern. Avoid deciding that a dog is well, unwell, happy or distressed from a single metric. Look at the dog's behavior and routine as a whole, and ask a veterinarian or qualified behavior professional when a change concerns you.
For practical household ideas, see our indoor dog enrichment guide. That guide has a different job from this Methods Note. It should not be used to turn the shelter study into a universal prescription.
Bottom line
This four-week study shows what careful measurement can add and what it cannot remove. Telemetry produced a detailed record of physiological patterns. The researchers then separated predictability, sleep quality and day-night cyclicity instead of treating them as one simple score.
The result is useful precisely because it stays narrow. Social conditions produced clearer changes in two measures for these eight dogs, while individual responses varied. A changed physiological pattern is not the same as a known emotional experience, and a study metric is not a household welfare guarantee.
Sources and update note
The primary study, its full text and the methods companion were re-opened on September 4, 2026. This is a bounded Methods Note, not a systematic review or meta-analysis. The study's measures and results are reported with the population, time window and comparison. Our non-inference boundaries are identified where the article explains what the sources cannot establish.
- Travain, T., Lazebnik, T., Zamansky, A. et al. "Environmental enrichments and data-driven welfare indicators for sheltered dogs using telemetric physiological measures and signal processing." Scientific Reports 14, 3346 (2024). Nature record · PubMed record · PMC full text
- Cobb, M. L., Jiménez, A. G. and Dreschel, N. A. "Beyond Cortisol! Physiological Indicators of Welfare for Dogs: Deficits, Misunderstandings and Opportunities." Journal of Applied Animal Welfare Science (2025). Formal DOI record · PubMed record
This article is informational and does not provide veterinary diagnosis or treatment advice.