Skip to content

fix(host-core): keep intact history readable around torn UTF-8 records - #1566

Closed
AR307 wants to merge 1 commit into
vastsa:mainfrom
AR307:codex/pr-transcript-utf8-followup
Closed

AR307 wants to merge 1 commit into
vastsa:mainfrom
AR307:codex/pr-transcript-utf8-followup

Conversation

@AR307

@AR307 AR307 commented Oct 11, 2026

Copy link
Copy Markdown
Contributor

Summary

Follow up on merged #1565 with a focused host-side fix for the UTF-8 failure discussed
in #1560. The outbox change prevents one refused append from blocking the
queue; this change keeps intact transcript records readable when an interrupted
append leaves an incomplete multibyte character. It is independently mergeable
and does not alter the outbox retry or quarantine policy.

Reproduced failure

The transcript readers use BufRead::read_line(&mut String). An invalid UTF-8
sequence fails during that read, before the existing malformed-record handling
can run. A torn Chinese/emoji record can therefore hide otherwise intact
history, including valid records appended after the damaged line.

All five added regression tests fail on unmodified main 22560e083aa7 with
stream did not contain valid UTF-8, and pass with this change. This establishes
the torn-record defect; it does not establish every reported occurrence's
origin or prove the proposed connection with #1515. ends_with_newline already
reads a raw byte and is left unchanged.

Changes

  • Find JSONL boundaries using bytes, then decode each record strictly.
    Log invalid UTF-8 without logging its content and skip that record on reads.
  • Advance offsets over complete invalid lines, but leave an unterminated tail
    eligible for refresh when the writer supplies more bytes.
  • Apply the same boundary to full, paged, indexed, tool-lookup and revision reads.
  • Preserve unrelated raw bytes and CRLF during metadata updates instead of
    requiring the entire file to decode as UTF-8.
  • Add focused regressions and extend the existing real-Electron regenerate/quit
    scenario to reopen a torn Unicode transcript before regenerating.

Reading does not truncate the transcript, replace characters, reconstruct lost
records or change the storage format. Invalid records stay on disk but are not
displayed. Existing explicit full-history rewrite behavior is unchanged.

Validation

Windows; existing Node 24 and Rust 1.90 toolchains; isolated app data and a
controlled local HTTP/SSE provider, without production credentials.

  • Rust transcript suite: 42 passed, including the five new Unicode tests.
  • Regression proof: the same five new tests against original main: 5 failed
    with invalid-UTF-8 errors, not fixture/build failures.
  • Adjacent Desktop persistence, regenerate, shutdown and subagent-settlement
    regressions: 28 passed, including the six new fix(desktop): quarantine outbox appends the host keeps refusing #1565 cases.
  • Workspace package builds, Desktop/pi-host typechecks, Electron
    main/preload/renderer build, fresh sidecar bundle and release Rust host build:
    passed.
  • Source lint, style-token and agent-policy checks, cargo fmt --check,
    cargo clippy -p host-core --all-targets --release --locked, and diff checks:
    passed.
  • node scripts/e2e-regenerate-quit.mjs --torn-unicode: 8 passed in the
    actual isolated Electron application. This covers a real Read tool response,
    reopening the same damaged session without a fork/model replay/file mutation,
    regenerate, native quit, restart and both revision branches. The resulting
    history screenshot was also inspected.
  • Full Rust comparison: candidate 809 passed, 4 failed, 1 ignored; unmodified
    main 804 passed, the same 4 failed, 1 ignored. Failures are the two existing
    config-sync engine cases and the two data-relocation capability-override
    cases. The full suite is not reported as green.
  • Rebased on main 06c04e99dc84 after fix(desktop): quarantine outbox appends the host keeps refusing #1565 merged and reran the combined
    Desktop build, adjacent regressions and actual Electron scenario.

macOS/Linux execution and the full Desktop/runtime test suites were not rerun
for this focused Rust follow-up.

Related to #1560; complements #1565.

Read JSONL boundaries as bytes before decoding each record so an
interrupted multibyte character cannot poison the whole transcript or
revision scan. Keep incomplete tails eligible for refresh and preserve
unrelated source bytes during metadata updates instead of truncating or
replacing them.

Cover Unicode boundaries, resumed appends, paged and indexed reads, tool
lookup and revisions. Extend the isolated full-desktop scenario to reopen
damaged history and regenerate across restart without replay on read.
@AR307 AR307 closed this Oct 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant