# AI chat portability: public-safe methodology and evidence notes

Prepared and reviewed by thredly, 6 September 2026.
Companion page: https://thredly.io/research/ai-chat-portability

This is a redacted editorial extract from first-party research records. It is not the original private transcript, a complete dataset, an independent audit or an original benchmark scorecard. Client labels A, B and C replace names. Square brackets replace identifying details; ellipses indicate shortened fragments. Monetary values and conditions necessary to explain the findings are retained. No additional research runs were conducted for this appendix.

## Scope

The supplied constructed stress conversation contains 323,457 Unicode characters and 2,902 literal speaker markers. It mixes current and superseded decisions, dates, tentative and signed terms, completed work, unresolved questions, rejected approaches and unrelated discussion. It is not verified natural customer history or a representative sample.

The reviewed source attachment exactly matches the current portability fixture. The historical architecture record identifies its benchmark case by size, not by a per-run input hash. Historical request identity therefore cannot be independently reconstructed from that record alone.

## A. Historical thredly handover benchmark

Source: internal production handover architecture lock, dated 22 August 2026. Relevant portions: Production decision, Locked architecture, Request path, Standard strict-entailment mode, Evaluation boundary and Known production limitation.

Minimal score excerpt: “User-supplied scores of 84 and 87, accepted mean 85.5; single call”. Markdown emphasis removed from this excerpt, wording unchanged.

The record identifies the selected production architecture as single-call Gemini 3.7 Flash (`gemini-3.7-flash`). It specifies full-source LF/NFC normalization, HIGH thinking, 32,768 maximum output tokens, zero automatic retries at launch, and final Markdown retained unchanged. No chunk-summary merge, extracted ledger, additional verifier/repair or automated scoring call forms part of that selected path. Offline gold-state evaluators and scorecards must not enter the production generation prompt.

The strict-entailment instructions require direct source support, respect user-authoritative corrections, preserve date/entity/scope and proposed/agreed/completed/rejected/unresolved distinctions, and avoid unsupported actions or generalized rules.

The accepted mean is (84 + 87) / 2 = 85.5. The attachment does not include original scorecards, category weights, a full rubric, scorer identity/process, run IDs or an explicit /100 denominator. This is a reported historical handover-quality result, not 85.5% accuracy or fact retention. The two original scored artifacts are not supplied in this evidence extract. The record describes a dated architecture decision, not independently verified execution traces for each scored run or today's production deployment.

## B. Separate full-transcript receiving test

Source: saved receiving-test input, staged-text inspection, raw completed answer and frozen current-state reference from 6 September 2026.

- Fresh Claude consumer incognito chat; displayed model Sonnet 5 Medium, Free plan.
- Complete 323,457-character source plus eight receiving questions: 327,331 characters in the staged input.
- The pasted-content preview contained that input exactly. This verifies staged UI text, not internal model attention or native chat restoration.
- One completed answer, no regeneration or silent correction; requested maximum 2,500 words.
- Reference and questions frozen before the test. The reference contained 65 entries supported by 353 distinct source turns, separating current agreement, tentative/internal state, rejected/superseded information, completed work and unresolved matters.
- The reference was not supplied to the receiving chat. Its construction was informed by earlier exploratory work; assessment was unblinded and not independently adjudicated.
- Assessment recorded retained information, omissions, stale/superseded state, unsupported interpretations and correctly identified uncertainty. Multiple labels can concern different subfacts. Irrelevant or closed history need not be repeated. No aggregate score or winner was produced.

The eight questions covered current and signed-future commercial terms; completed customer work; employment/team/compensation state; vehicle commitments; household administration; travel/claims/events; unresolved work and exclusions; and which source materials actually arrived. They required the final narrative state, explicit uncertainty, no invented facts and no external actions. This is a topic summary, not a reproduction of private names in the prompt.

## Minimal redacted evidence fragments

### B1. Client A: old price attached to current dates

Earlier source: “[Client A] … flat £18k annual contract ending [earlier end date].”

Later source: “[Client A] signed. 24 months from [start date] to [current contract end], £19,000 per year flat fee …”

Final correction: “Current £19k contract still runs through [current contract end], so don't start 19.5 early.”

Receiving answer: “£18,000/year flat, ~[user count] users, SAML only, runs to [current contract end]”.

Finding: the answer combined a superseded price with the current end date. Its signed-future renewal row separately preserved the future £19,500/year agreement and main terms correctly. Correct future terms do not repair the wrong current rate.

### B2. Client B: request-conditioned export omitted

Source condition: “within 10 business days of written request made within 30 days after expiry or termination.”

The final agreement continued other terms. The answer retained the breach-cure period but omitted this export request condition. This is a reconstruction omission, not a failed live export/import test.

### B3. Client C: a superseded price carried into renewal discussion

Earlier internal proposal: “£175/month from [proposed start] … We haven't offered it yet …”.

Later source: “let free SCIM continue through current contract end [date] as a goodwill bridge, then price everything in renewal from [renewal start]. That's approved internally but not told to [Client C] yet.”

Receiving answer: “then price at renewal (£175/mo considered)”.

Finding: the later source replaced the earlier short amendment approach; it did not approve that amount for the renewal. The answer carried the obsolete proposal into its renewal discussion, but did not claim the customer had signed or agreed that price. It correctly retained the unresolved customer-facing status.

### B4. Correct reconstructions

The answer distinguished an achieved personal bonus component from an unresolved overall payout. It also kept completed house-move proceeds separate from later improvements. No identifying personal amounts, names or documents are included here.

## Interpretation and limitations

A evaluates a generated handover using historical external scoring. B evaluates a fresh chat's answers using a frozen item-level reference. Different tasks, prompts, answer budgets and evaluation methods mean no controlled head-to-head or comparable score exists. The historical 85.5 must not be attributed to the new study or any separate later handover.

One case, one receiving answer, incomplete historical score provenance and unblinded assessment limit inference. The answer budget may affect omissions. A staged input match cannot prove complete model attention. Native account-wide export/import and ThreadPort routes were not completed; access/privacy restrictions are not provider failures.

This case shows that intact text transport did not guarantee correct reconstruction of current working state. It does not show that thredly beat Claude, that full transcripts are always worse, that a handover is lossless or that these findings represent all users/models. A structured handover may help expose current state explicitly, but improved downstream performance is not established by these separate evaluations.

Research and product-interest disclosure: thredly produced this analysis and offers a structured conversation handover product. This is not an independent product review. The complete private transcript, account screenshots and private research ZIP are deliberately not included.
