Similarity Measures for Small Data Problems

When you can safely reuse someone else's data

This protocol makes an analysis team say which similarities their result depends on before they combine two datasets, and which of those they can actually check.

Period
2025–present
Role
Lead author
Stack
Evidence synthesis, Meta-analysis, Data integration

Problem

Combining data across sources is routine. Teams pool two cohorts, reuse a published model at a new hospital, borrow controls from an old trial, or substitute observational data where a randomized trial is not available. Every one of those moves rests on a claim that the sources are similar enough for the specific job at hand. What usually gets checked instead is whatever happens to be convenient to measure, most often whether the covariate distributions overlap. The dissimilarity that would actually break the result goes unexamined, and nothing in the workflow flags its absence.

Approach

The similarity chain reverses the order. The analytical task is fixed first, and it determines what has to be compared. Each link ties the task to the kind of similarity it depends on: of the data, of the design that produced them, or of the model being reused. The link then states a plain-language condition that similarity has to meet, and an empirical check for the condition. Some conditions act as gates: when a gate fails, everything downstream is undefined rather than merely imprecise, and the chain says so rather than returning a number. The paper works the protocol through a published meta-analysis, where the pooling decision turns out to rest on an assumption that the documentation supports but no measurement can confirm.

Outcome

The protocol ships as working material: a checklist, reference catalogs that map common tasks to the similarity types they require, and a conversational agent skill that walks a researcher through the specification and writes out a report for a methods section. The paper carries further worked examples besides that one. They cover pooling external data for a target setting too small to support a model on its own, substituting real-world evidence for a randomized trial, and cross-species single-cell integration.