Executive answer
Moving every word is not the same as preserving the current working state.
In one 323,457-character stress case, a fresh Claude chat supplied with the full transcript still combined an obsolete price with current contract dates and omitted some requested conditions. Other details were reconstructed correctly.
Separately, the historical thredly production record reports handover-quality scores of 84 and 87, with an accepted mean of 85.5 for its 323k stress case. The evaluations used different methods. They do not establish that thredly beat Claude.
Read the evidence boundaries ↗01 / The distinction
Two different jobs, often mistaken for one
Text transport means moving the intended source material into a new chat. Current working state means the decisions, constraints, relevant completed work and unresolved tasks that remain authoritative at the end of that conversation.
Did the material arrive?
The full history includes old offers, corrections, abandoned ideas and the final decision. All can be present at once.
Which parts are true now?
The receiving chat must resolve what changed, preserve conditions and avoid turning proposals or completed tasks into live commitments.
A current-state handover is a structured account of that authoritative state for a fresh conversation. It is a selection and reconciliation task, not simply a shorter transcript. It can also contain omissions or mistakes.
02 / The case
Long enough for yesterday’s answer to become today’s error
The supplied stress conversation contains 323,457 characters and 2,902 literal speaker markers. It was deliberately constructed to mix changing commercial terms, personal administration and unrelated conversation. It is test material, not verified natural customer history or a representative sample of AI users.
- Current decisionsAlongside the decisions they replaced.
- Changing dates and pricesWith corrections that apply to a specific agreement.
- Tentative and signed termsInternal approval is not customer agreement.
- Completed and unresolved workA closed task should not become a new to-do.
- Rejected approaches and noiseIncluding explicit instructions not to carry details forward.
The difficulty is not just length. It is keeping each fact attached to the right entity, period and status. A plausible number paired with the wrong date can look convincing while being materially wrong.
03 / Full-transcript findings
The whole transcript arrived. An old price survived.
The complete source and the receiving questions were found intact in Claude’s staged pasted-content preview. The resulting answer nevertheless made the following reconstruction errors. Client labels below are anonymised; fragments are shortened and redacted. Evidence B
£18,000 per year
Earlier contract period
£19,000 per year
Current contract period
Receiving answer · redacted“£18,000/year flat … runs to [current contract end]”
The answer attached the superseded price to the current agreement. The source repeated the correct £19,000 rate near the end.
| Case | What the source established | What the answer did |
|---|---|---|
| Client B Export conditions | Export within 10 business days of a written request made within 30 days after expiry or termination. | Omitted the request conditions. This was a missing contractual detail, not a test of an export feature. |
| Client C Superseded pricing | An internal £175/month proposal was superseded by a free bridge and unresolved renewal pricing. | Wrote “then price at renewal (£175/mo considered)”. This carried the old proposal into renewal discussion. It did not assert a signed price. |
What the answer got right
Reconstruction was not uniformly wrong
- It separated Client A’s signed future renewal from the active agreement and retained the main future terms.
- It distinguished an achieved personal bonus component from an unresolved overall payout.
- It preserved the boundary between completed house-move proceeds and later improvement spending.
The requested answer ceiling was 2,500 words, which may contribute to omissions. The evidence does not isolate why the stale price was selected, and does not show that the model attended to every passage.
04 / Separate historical evidence
What the thredly benchmark actually recorded
The production architecture record locked on 22 August 2026 reports two externally supplied handover-quality evaluations for its 323,457-character stress case. It identifies the selected production architecture as single-call Gemini 3.7 Flash. Evidence A
These are reported scores, not accuracy percentages. The supplied record does not include the original scorecards, detailed rubric or an explicit denominator. We do not independently establish a “/100” scale.
The documented workflow sends the complete normalized conversation to Gemini once, with strict-entailment instructions to preserve source-supported facts, corrections, scope and state. It returns the final Markdown handover directly. There is no chunk-summary merge or additional repair or scoring call in that architecture. Gold-state evaluation is kept outside the generation prompt.
These historical evaluations concern the handover itself. They did not measure answers to this portability study’s receiving questions. The accepted mean is not a new portability-study score and must not be assigned to a separate later handover.
05 / Methodology
One source. Two distinct evaluations.
| Method | A · Historical handover benchmark | B · Full-transcript portability test |
|---|---|---|
| Evidence date | Record locked 22 August 2026 | Receiving test completed 6 September 2026 |
| Task | Produce a continuation-ready current-state handover. | Answer eight questions about current state using the complete transcript. |
| Model/workflow | Recorded single-call Gemini 3.7 Flash; strict-entailment instructions. | Fresh Claude consumer incognito chat; displayed Sonnet 5 Medium, Free plan. |
| Evaluation | Two externally supplied scores accepted in the production record. Original rubric not supplied. | One answer assessed against 65 frozen reference entries, supported by 353 cited source turns. No aggregate score. |
| What is inspectable | Dated architecture record and reported results; not the original scorecards. | Saved input, staged-text check, raw answer and item-level assessment. Public fragments are redacted. |
How the receiving answer was assessed
The current-state reference and eight questions were frozen before the receiving test. The source-backed reference separated current agreements, tentative/internal plans, superseded state, completed work and unresolved matters. It was not supplied to the receiving chat.
One completed answer was retained unchanged, without regeneration. Assessment recorded correctly retained details, omissions, stale/superseded facts, unsupported interpretations and correctly identified uncertainty. Omitting irrelevant or closed history was not automatically treated as harmful loss.
The reference was informed by earlier exploratory work. Assessment was unblinded and not independently adjudicated. Literal role labels in a pasted transcript are not native chat turns.
06 / Interpretation
What this evidence does and does not show
It shows a concrete failure of current-state reconstruction
An intact staged transcript did not prevent the receiving answer from using a stale fact. It also records that an earlier thredly handover workflow received two accepted external evaluations on its 323k stress case.
It does not establish comparative superiority
- Different tasks, prompts, outputs and evaluation methods prevent a controlled head-to-head comparison.
- One constructed case and one new receiving answer do not represent all models, users or conversations.
- Historical scoring details are incomplete. The architecture record identifies its case by size; historical per-run input hashes and scorecards were not supplied.
- A staged-text match establishes UI transport, not complete model attention or native conversation restoration.
- No claim is made that full transcripts are always worse or that structured handovers preserve everything.
- Native account export/import and ThreadPort workflows were not completed. Access restrictions are not transfer failures.
- This is research produced by thredly about a problem its product addresses, not an independent product review.
The practical implication
Capacity to accept a conversation is not a guarantee of understanding its final state.
A structured handover may make the current decisions, constraints and unresolved work easier to carry into a fresh chat, even when the full transcript fits. That is a useful rationale to investigate, not a downstream performance advantage proven by these two evaluations. Review any generated handover before relying on it.
07 / Evidence notes
Sources, redactions and provenance
- A · Historical production architecture record. Internal first-party record, locked 22 August 2026. Reports scores 84 and 87, mean 85.5 and the selected single-call architecture. This is attribution to a dated record, not independent re-scoring or a verification of today’s live deployment.
- B · Full-transcript test record. Existing 6 September 2026 run: exact staged input check, unchanged receiving answer, frozen 65-entry current-state reference and item-level assessment. The reviewed attachment exactly matches the current study source.
- Public evidence handling. Client names are replaced with A, B and C; identifying dates and unrelated personal details are omitted. Ellipses and square brackets mark shortened or redacted fragments. The full private transcript, account screenshots and research ZIP are not published.
Download the public-safe methodology and evidence appendix (Markdown) ↓
The appendix contains only a redacted methodology summary and minimal supporting fragments. It is an editorial extract, not the original private record, a complete dataset or an independently audited scorecard.