AN INQUIRY INTO UNPAIRED ROSETTA 10 OCT 2026

Do the shadows agree,
or the things
themselves?

Let us ask of Unpaired Rosetta: how much of the agreement belongs to what is seen, and how much to the models through which we see it?

WHAT THE TRIAL ESTABLISHES

We have neither replicated
nor disproven Rosetta.

Our simpler distance-matching test failed: none of six model pairings recovered more than one of the five human labels.

The five-photo trial did not run Rosetta’s alignment method. The next step is a direct test.

THE WALL AND ITS SHADOWSFIG. 01
The wall and its shadows An imagined scene from Plato's cave. Fire behind a low wall casts the shapes of carried objects onto stone. A seated figure watches the shadows. A narrow opening leads toward daylight.

This scene is an allegory made for the inquiry. The measured evidence is given below.

Six model pairings · comparison completeThe same five photographs · every result retainedGloVe: 1/5 each · ELMo: 0/5 each

THE SAME FIVE PHOTOGRAPHS · SIX MODEL PAIRINGS

Change the instruments. Keep the things.

The models changed.
The names still went astray.

We kept the five photographs, human labels, crop and distance rule, and compared three image models with two language models. None of the six pairings recovered more than one correct match. Here is the whole account.

THE VERDICT ON THIS TRIAL

Our simpler method failed on these five photographs. Rosetta has neither been replicated nor disproven by this trial. We matched the distances among five items directly. Rosetta instead learns a map through cluster matching and refinement, then evaluates it on unseen examples. Read the authors’ method.

What each pair chose.

With GloVe, each image model matched one name correctly. With ELMo, none did. DINO and MoCo chose the same assignments for each language model. MAE changed the assignments, although its number of correct matches stayed the same.

Correct matches out of five, followed by the human assignment’s rank among all 120 assignments. A smaller rank means that assignment has a smaller distance error. The original DINO/GloVe features and result are reused unchanged.
Image modelGloVeELMo
DINOv11/5Human rank 94/1200/5Human rank 26/120
MoCo-v21/5Human rank 91/1200/5Human rank 40/120
MAE1/5Human rank 99/1200/5Human rank 24/120

All six optima were unique at the fixed tolerance of 10−12. These are six views of the same five photographs, not six independent samples. A random assignment gets one name right on average. The table supplies no significance test or estimate of accuracy on new photographs.

WHAT CHANGED, AND WHEN

Six comparisons. Still five observations.

This follow-up was designed after the first result. Its added models came from the earlier provenance search; its code and rules were published before their features and scores were computed. Each pairing examines all 120 assignments without receiving the correct pairs. There is no unseen evaluation set, and this is not a run of Rosetta.

The added MoCo, MAE and ELMo files were downloaded and hashed after the collector reports taking the photographs. Their documented training histories are older, but they do not share the original pair’s file lock before capture. That is a weaker chronology. Neither history proves that familiar objects or places were absent from training.

See every chosen name
Each row names a photograph by its human label. The remaining cells show the label each model pair assigned to it. The scoring step used the human correspondence after all six selections were saved. The follow-up was designed after the first result; photographs and raw features remain private.
Photograph’s human labelDINO + GloVeDINO + ELMoMoCo + GloVeMoCo + ELMoMAE + GloVeMAE + ELMo
SlideBasketballTreeBasketballTreeTrashcanBasketball
TrashcanTreeWater fountainTreeWater fountainTreeWater fountain
TreeTrashcanBasketballTrashcanBasketballBasketballSlide
BasketballSlideSlideSlideSlideSlideTrashcan
Water fountainWater fountainTrashcanWater fountainTrashcanWater fountainTree
Which model outputs were compared?

The image features are the CLS output of original DINOv1 ViT-B/16, the pooled backbone of original MoCo-v2 ResNet-50, and the CLS output of original MAE ViT-B/16. The language features are GloVe 6B, 300 dimensions, averaged over label tokens; and original ELMo trained on the 1 Billion Word Benchmark, averaging the top contextual layer over real tokens with fresh recurrent state for each label. Every vector is normalized to unit length.

On these five photographs, MAE’s cosine distances range from 0.0054–0.0397, compared with 0.4923–0.9231 for DINO and 0.3655–0.8125 for MoCo. The distance scales differ, so raw matching costs must not be used to rank the models. These one- or two-word labels also say little about what ELMo would do with richer descriptions.

What the comparison can tell us. Differences between cells show how this fixed distance rule responds to these model choices on these five photographs. Models also differ in training data, architecture and objective; these comparisons cannot separate those causes.

What remains to be asked. A collection made after the whole panel is fixed would give a stronger chronology. New scenes would be needed to ask whether the result travels beyond this outing. Neither this table nor the original trial settles the larger claim.

Run on 10 October 2026. Code and rules were published before the added feature extraction. Independent arithmetic replay agreed with all 720 assignments. It used saved features; it did not rerun the encoders or authenticate capture time or training history. The first trial below remains the original reference; the two earlier collection plans remain separate records.

01 / THE QUESTION

What follows from agreement?

When two shadows agree,
what have we learned?

The chosen encoders may share training data, objectives, or useful biases. If we repeat the solver but keep the encoders, we have not yet asked whether other models would agree.

A

Would other models agree?

A new alignment seed gives another trial with the same embeddings. The completed comparison changes the image and language models, keeps the five photographs, and reports every pairing. It was designed after the first result.

Six pairings on five photographs
B

Have they seen the same things?

Two datasets may bear different names and still contain the same images. The metadata shows that some COCO source images are described in Stanford Paragraph Captioning (SPC).

Metadata examined
C

The thing, or merely its kind?

To find a horse is not yet to distinguish this horse, its number, its place, or what it does. Retrieval must also be tested among things of the same kind, where broad categories no longer suffice.

Further trials proposed

THE ORIGINAL FIVE-PHOTOGRAPH TRIAL · REFERENCE

Five things. Five names.

Can their distances
recover their names?

Slide, Trashcan, Tree, Basketball and Water fountain: the five human labels, each supplied with a new photograph. We asked whether the arrangement of the image features resembled the arrangement of the word features closely enough to recover their correspondence.

THE QUESTION WE ACTUALLY TESTED

Which of the 120 assignments makes the distances agree?

Original DINOv1 describes the photographs; original GloVe describes the exact human labels. Compare the ten distances between image pairs with the ten distances between word pairs. Examine every one-to-one assignment and choose the smallest mean squared difference. The matcher receives no correct image–label pairs.

We chose this test after seeing the photographs, then published its code and rules before computing their features or scores. All five images and the complete set of candidate labels take part in selection. This is a small exploratory test, with no unseen evaluation set.

The assignment it chose.

One correct match out of five.

Photographs are named below by the human label for evaluation. Those names were withheld from the matcher. The photographs themselves remain private.
Photograph’s human labelChosen labelMatch
SlideBasketballNo
TrashcanTreeNo
TreeTrashcanNo
BasketballSlideNo
Water fountainWater fountainYes

The human assignment ranked 94th of 120 by the fixed distance rule. The selected assignment had cost 0.037819; the human assignment, 0.073473. The optimum was unique at the prescribed tolerance of 10−12, with a gap of 0.000281 to the next cost. Cost is the mean squared difference between ten pairs of distances.

For the selected assignment, those distance pairs had Pearson correlation 0.878. Four names still went to the wrong photographs. One correct match is also the average under a random assignment; this single result does not establish a general level of accuracy.

I

The scene came later.

The collector attests that all five photographs were taken after the model files were fixed. If that account is true, these particular captures could not have trained those files. The attestation is not independent proof of capture time; familiar objects, words and places remain possible training material.

Capture time attested by collector
II

Let the result keep its size.

Five photographs from one outing are five observations. The 120 assignments are alternatives, not 120 independent trials. A random assignment gets one of five right on average. We report no significance test or general accuracy estimate.

One chosen model pair
III

Do not repair the answer afterward.

Keep the supplied words, fixed center crop and distance rule. Shared backgrounds, ambiguous words and the crop may affect the result. This calculation trains no cross-modal map and is not a run of the Rosetta algorithm.

Original model pair only

Run on 10 October 2026. The distance objective failed to recover four of the five human matches in this set. This does not refute Rosetta or predict results on new photographs. The comparison above records six model pairings on this same set. Code and protocol were published before inference. The two earlier proposed collections below remain separate records.

THE EARLIER 48-PHOTOGRAPH PLAN · NOT RUN

Keep the proposal distinct from the trial.

The larger collection.
A plan kept apart.

This earlier proposal asked whether two frozen models support correspondence between new photographs and descriptions when supplied with correct fitting pairs. Its 48 photographs have not been collected. The five-photo trial above follows different rules and does not complete this plan.

THE REASON FOR THAT PROPOSAL

The scene comes after the model.

Publish the exact model hashes, collection rules and analysis code before taking the photographs. If that order is honestly followed, these particular recorded events cannot have trained those already fixed files. We need not rely only on the model publisher’s account of its old training corpus.

This is a narrower defence, not a claim of complete purity. The models have seen mugs and books; ordinary phrases may recur. Markers and file hashes support the collection record but cannot independently prove capture time. The choice of models remains an open question.

  1. Before

    Fix what will see and read.

    Original DINOv1 ViT-B/16 and GloVe 6B/300d. Publish their byte hashes, image crop, text rule, sample and six fits. Recheck the files before inference. Read the fixed protocol.

  2. 48 photographs

    Bring new scenes before them.

    Eight kinds of everyday object, photographed in six rounds. Four fitting rounds use one set of physical objects; two development rounds use a different set. A person writes the captions after all photographs exist.

  3. After

    Ask whether a known pair can meet.

    Run separate image and word noun probes, a paired ridge map and three maps fitted with shuffled pairs. Keep repeated descriptions, exact ties and failures in the account. Balanced random choice is 12.5%.

I

A small trial, with a small claim.

Mug, bowl, plate, bottle, spoon, fork, book and shoe. Gather two physical examples of each: sixteen objects. The sixteen development photographs show eight held-out objects twice. They are not sixteen independent instances.

32 fit · 16 development
II

First test whether the features suffice.

The correct pairs would be supplied under this proposal. If even that mapping failed, an unpaired failure would be hard to interpret. A useful paired result would give reason to design another trial; it would not establish unpaired alignment.

Paired competence fits proposed
III

Leave room for the answer.

No tuning, added photographs or changed categories after seeing scores. No final test has been collected. Two development rounds and three shuffle references cannot bear a claim of statistical significance.

This proposed study has not run

Earlier prospective protocol · 10 October 2026. Preserved collection record · The 48-photograph collection has not begun. Software checks used artificial fixtures. This plan and the paused EPIC plan below remain separate from the five-photo trial.

THE EARLIER EPIC PLAN · PAUSED

Let the claim be no larger than the trial.

The earlier plan.
Its record remains.

Can one fixed image model and one fixed word model recover object-category correspondence on later kitchen recordings, when their alignment receives no paired episodes? That was the proposed EPIC test. Its protocol and limits remain available; the five-photo trial does not resume or complete it.

WHY BEGIN HERE?

A narrower claim is easier to defend.

The training corpora predate these recorded episodes. Under the published training histories, memorizing these particular episodes cannot explain a result. This makes the case stronger against direct episode exposure. Familiar objects, ordinary phrases and the choices of the model builders remain part of the account.

The question was narrow; the proposed run was substantial. One model pair, one source of recordings and one primary measure keep the claim narrow. The frozen plan calls for 9,392 clips, 37,568 frames and 17 fits. Those alignment fits have not run.

  1. 2012

    The images

    The ImageNet-1K collection, used without class labels by original DINOv1 ViT-B/16. DINO training record · ImageNet release

  2. 2014 and earlier

    The words

    Original GloVe 6B: Wikipedia 2014 and the older Gigaword 5 collection. We average its word vectors without further language training. Original GloVe release

  3. 2017

    The later episodes

    EPIC-KITCHENS-55 recordings and their human narrations. The recording years come from the per-video metadata. Recording dates

These are corpus and recording dates, not model publication dates. The argument depends on the documented training histories and publisher-supplied dates. A short phrase may occur in both old text and a later narration; that alone is not exposure to the later episode.

I

Keep the test apart.

The existing sample was fixed before alignment scores: 6,639 fitting clips, 1,127 development clips and 1,626 test clips. The seven test kitchens appear in neither other partition.

Selection checked
II

Withhold the correspondence.

The main fit receives images from 73 videos and text from 74 different videos. Known-pair mappings, shuffled-pair mappings, random maps and separate encoder probes help us interpret success and failure.

17 fits prescribed · paused
III

State what a match means.

The first measure asks whether the nearest description has the same annotated noun class. Each test kitchen receives equal weight; duplicate descriptions and exact ties are accounted for. Recognizing a cup does not yet distinguish this cup or what is done with it.

No benchmark scores

What could follow from the result?

Possible outcomes, not findings. Negative controls use shuffled pairs or random maps. Read differences with the prescribed uncertainty intervals and results for every seed.
What we observeWhat we may conclude
Unpaired retrieval exceeds the negative controls; paired retrieval also works.Evidence of recoverable object-category correspondence for this pair on later episodes. Shared concepts and model biases remain possible explanations for that correspondence.
Paired retrieval works; unpaired retrieval does not.The frozen features support a useful mapping with pairing information. The tested unpaired procedure has not recovered it under these conditions.
Neither paired nor unpaired retrieval works.The features, pooling, task or domain may be unsuitable. This outcome alone cannot identify the unpaired solver as the cause.

Unexpected behavior between controls also needs investigation. No outcome from this one pair establishes a universal representation, understanding of actions or syntax, or independence from model choice. Changing both the models and the benchmark cannot measure how much contamination affected the original COCO/SPC result.

The run is paused. Acquisition and its supervisor were suspended on 10 October 2026 before any full EPIC alignment scores. Existing data, caches and the fixed protocol are preserved. The five-photo trial and the unrun 48-photo proposal are separate records.

The controls have limits too. Paired and shared-episode controls use all fitting clips, a larger information budget than the main disjoint fit. Annotated action windows supply preprocessing supervision. These comparisons cannot isolate contamination by themselves.

Work record · 10 October 2026. Paused-run record · Code and checks. The earlier pilot below used different models and data.

EXAMINING THE SOURCES

Two names.
Some of the same images.

We joined the SPC, Visual Genome, and COCO metadata. In the released cross-dataset setup, some descriptions and images refer to the same source.

Read the source record
6,258unique COCO training images
described in SPC
3,335unique COCO validation images
described in SPC

This establishes shared sources. It does not show that the aligner was told which items were pairs. The effect on published performance is still unknown; this finding does not concern the separate disjoint-half COCO experiment.

02 / THE EARLIER PILOT

The first check held the models still.

Five conditions.
The same two models.

Keep DINOv2-B and MPNet fixed. Remove the known shared sources, then compare each removal with a random-removal control of the same size.

O / ORIGINAL

Begin with the original SPC population.

Keep the known validation and training source matches. This gives us the starting condition.

WHAT THIS COMPARISON ASKS

What do we observe before removing either kind of shared source?

19,561SPC paragraph rows retained
Starting condition
0Original population: 19,561 rows
Known validation source
3,337
Known training source
6,261
No known COCO match
9,963
Removed
0

These are the full populations, before the pilot’s 4,096-row cap. Here we count paragraph rows; above we count unique COCO images. Random controls use selection seed 0.

Ask each the same questions. Every arm uses the same three 2,048-query sets and complete 40,504-item gallery. Clean, all-population, and originally exposed queries are sampled separately.

Say only what was removed. “Source-disjoint” means known metadata identities only. Missing mappings, visual near-duplicates, and pretraining overlap remain unresolved. Random deletion matches size, not semantic composition.

03 / WHAT HAS BEEN SHOWN

Recorded · 10 October 2026

The trials ran.
What did they establish?

Each of the five conditions completed three alignment seeds. The reduced pilot shows that the runs finish, the exclusions are applied, and the saved results pass consistency checks.

Each fit uses 4,096 training rows per modality. These equal caps remove the full-population size contrasts. The scores therefore do not estimate the effects of the planned full-data exclusions.

Read the checks
15 / 15fits completed
92,160query records verified across 45 strata
  • Fixed query and gallery identities checked
  • Source-identity exclusions checked
  • Saved map, rank, and configuration hashes checked
  • Mean rank, median rank, and FOSCTTM recomputed from saved ranks

Verification did not recompute embedding similarities or Recall@k from scratch.

Examine all 15 reduced pilot runs

FOSCTTM below uses the same clean queries against the full gallery. Lower is better. These reduced runs check the procedure; their scores do not answer the scientific question. Every seed is shown.

Reduced pilot runs · DINOv2-B × MPNet
ArmFit seedStatusFOSCTTMFit time
Download all pilot measurements (JSON) ↓

THE NEXT STEPS · PROPOSED

First put their claim to their test.

Now test Rosetta
by its own method.

The five-photo comparison is complete. It tested a simpler rule. The earlier reduced runs checked the software; neither experiment reproduces a full published result. The immediate task is one faithful baseline, before another collection or a larger model search.

  1. 01

    Reproduce one published result.

    Use the authors’ code and released embeddings with one documented model pair, the original split, full settings and evaluation measure. Fix the seeds and comparison rule before running. Compare against the corresponding published result. If it does not reproduce, investigate the implementation and conditions before interpreting new failures. Released embeddings check the alignment procedure; they do not independently verify encoder training or extraction. The authors’ code and instructions.

    FIRST TASK · NOT RUN
  2. 02

    Change one condition at a time.

    Keep the data, split and Rosetta procedure fixed while replacing one encoder at a time. The paper already examines several model combinations; choose and justify additions before seeing scores. Include a mapping learned from known pairs and shuffled-pair controls. If the paired mapping works but Rosetta does not, that narrows the failure to recovering a useful map without the correct pairs under these conditions. Test known source overlap separately, with unchanged encoders and size-matched random removal.

    AFTER THE BASELINE
  3. 03

    Collect after the whole panel is fixed.

    Freeze every model file, readout and rule before collecting new photographs and human descriptions. Choose a justified sample size with enough fitting data, separate development and final evaluation scenes, several sessions, and multiple objects of the same kind. This gives a stronger case against training on those exact later captures. It does not exclude familiar concepts or prove that the pretraining corpora were disjoint.

    A NEW PROTOCOL

These are proposed next steps, not a frozen protocol or newly launched run. Preserve all six completed results. No further photographs or training from scratch are needed to begin the first step.

THE RECORDS

Examine the account.

The protocol and evidence records are below. They describe the work as it stood on the date shown; they do not update as experiments run.