Would other models agree?
A new alignment seed gives another trial with the same embeddings. To speak of other model families, we must choose them before seeing the results, then report every pairing and failure.
Not yet testedAN INQUIRY INTO UNPAIRED ROSETTA 10 OCT 2026
Let us ask of Unpaired Rosetta: how much of the agreement belongs to what is seen, and how much to the models through which we see it?
This scene is an allegory made for the inquiry. The measured evidence is given below.
01 / THE QUESTION
What follows from agreement?
The chosen encoders may share training data, objectives, or useful biases. If we repeat the solver but keep the encoders, we have not yet asked whether other models would agree.
A new alignment seed gives another trial with the same embeddings. To speak of other model families, we must choose them before seeing the results, then report every pairing and failure.
Not yet testedTwo datasets may bear different names and still contain the same images. The metadata shows that some COCO source images are described in Stanford Paragraph Captioning (SPC).
Metadata examinedTo find a horse is not yet to distinguish this horse, its number, its place, or what it does. Retrieval must also be tested among things of the same kind, where broad categories no longer suffice.
Further trials proposedTHE PRESENT TRIAL
Let the test come after the teaching.
Original DINO learned from the ImageNet collection; original GloVe learned from older text. The EPIC-KITCHENS episodes were recorded later. If the published training histories hold, these particular recordings could not have taught either model.
ImageNet-1K, used without class labels by the original DINOv1 image encoder.
Original GloVe 6B: Wikipedia 2014 and the older Gigaword 5 news collection.
EPIC-KITCHENS-55: recordings of kitchen work, followed by human narrations.
These are corpus and recording dates, not model publication dates. The chronology excludes the particular episodes under the documented histories. Familiar objects, ordinary phrases, and human choices in collecting the training data remain.
The sample was fixed before scores: 6,639 fitting clips, 1,127 development clips and 1,626 test clips. The seven test kitchens appear in neither other partition.
Selection checkedThe main fit receives images from 73 videos and text from 74 different videos. Paired, shuffled and random-map controls ask what the same frozen features can support.
17 fits prescribedThe first measure asks whether the nearest description names the same object class. Each test kitchen receives equal weight. This does not establish motion, syntax, or the identity of a particular episode.
Scoring pendingThe work is in progress. Four frames per clip are being acquired: 37,568 in all. Missing or invalid examples stop the run. The recorded decoding failure and its bounded recovery check remain in the account.
A stronger control has a cost. Paired and shared-episode controls use all fitting clips, a larger information budget. Their comparison with the disjoint fit cannot isolate contamination. Annotated action windows also supply preprocessing supervision.
Work record · 10 October 2026. This page is a dated account, not a live monitor. The earlier pilot below used different models and data.
EXAMINING THE SOURCES
We joined the SPC, Visual Genome, and COCO metadata. In the released cross-dataset setup, some descriptions and images refer to the same source.
Read the source recordThis establishes shared sources. It does not show that the aligner was told which items were pairs. The effect on published performance is still unknown; this finding does not concern the separate disjoint-half COCO experiment.
02 / THE EARLIER PILOT
The first check held the models still.
Keep DINOv2-B and MPNet fixed. Remove the known shared sources, then compare each removal with a random-removal control of the same size.
O / ORIGINAL
Keep the known validation and training source matches. This gives us the starting condition.
What do we observe before removing either kind of shared source?
These are the full populations, before the pilot’s 4,096-row cap. Here we count paragraph rows; above we count unique COCO images. Random controls use selection seed 0.
Ask each the same questions. Every arm uses the same three 2,048-query sets and complete 40,504-item gallery. Clean, all-population, and originally exposed queries are sampled separately.
Say only what was removed. “Source-disjoint” means known metadata identities only. Missing mappings, visual near-duplicates, and pretraining overlap remain unresolved. Random deletion matches size, not semantic composition.
03 / WHAT HAS BEEN SHOWN
Recorded · 10 October 2026
Each of the five conditions completed three alignment seeds. The reduced pilot shows that the runs finish, the exclusions are applied, and the saved results pass consistency checks.
Each fit uses 4,096 training rows per modality. These equal caps remove the full-population size contrasts. The scores therefore do not estimate the effects of the planned full-data exclusions.
Read the checksVerification did not recompute embedding similarities or Recall@k from scratch.
FOSCTTM below uses the same clean queries against the full gallery. Lower is better. These reduced runs check the procedure; their scores do not answer the scientific question. Every seed is shown.
| Arm | Fit seed | Status | FOSCTTM | Fit time |
|---|
THE RECORDS
The protocol and evidence records are below. They describe the work as it stood on the date shown; they do not update as experiments run.