perf(web): highlight chat code off the main thread - #626
Open
Andrey Markin (Mark-Life) wants to merge 10 commits into
Open
Andrey Markin (Mark-Life) wants to merge 10 commits into
Andrey Markin (Mark-Life) wants to merge 10 commits into
Conversation
This was referenced Oct 4, 2026
Andrey Markin (Mark-Life)
marked this pull request as ready for review
October 4, 2026 17:55
Andrey Markin (Mark-Life)
requested review from
Olga Lavrichenko (OLavrik),
Rustam Sadykov (SBOne-Kenobi),
danyaberezun and
Rinat S (rsolmano)
as code owners
October 4, 2026 17:55
Andrey Markin (Mark-Life)
marked this pull request as draft
October 5, 2026 08:43
Andrey Markin (Mark-Life)
force-pushed
the
perf/shiki-worker
branch
from
October 5, 2026 08:47
1a4e7aa to
5b92aa5
Compare
Andrey Markin (Mark-Life)
marked this pull request as ready for review
October 5, 2026 09:27
Andrey Markin (Mark-Life)
force-pushed
the
perf/shiki-worker
branch
from
October 5, 2026 15:17
5b92aa5 to
b2bab48
Compare
Opt-in scenarios for perf:render, selected with --scenario: - long-stream: one ~26k-char seeded markdown answer in 10-60 char deltas - parallel-agents: 20 chat tabs streaming at once - background-agents: 19 hidden tabs streaming, visible chat idle --cpu N throttles the CPU through CDP during the measured window. Every run now also records long tasks, rAF frame gaps, CDP main-thread time, Markdown renders and self ms per delta, and scenario counters. The output records rootDir, dirty flag and a dirty-tree hash so runs from other worktrees stay comparable. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…scenario Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Assistant text re-parsed the whole message on every streamed delta, so cost grew with message length. BlockMarkdown splits the text into top-level blocks with the marked block lexer and renders each block with a memoized react-markdown, so a delta re-parses only the tail. The splitter is incremental: it re-lexes from the last closed block that has two blocks after it. A block closes at a blank line, or at a heading, a rule or a CommonMark-closed fence, so text without blank lines still splits. A link or footnote definition, also inside a list item, renders the message whole. Every fallback (open HTML, raw HTML lines, lazy list lines, indented code, a fence only marked closes, lossy lexer output) merges blocks and never splits them. Blocks are separated by the same newline the whole-document render emits, so the HTML is identical. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Check blank lines against the previous token's trailing whitespace, so a long open block splits in linear time. - Let a closed HTML block followed by a blank line end its block, so back-to-back <details> sections split. - Count only ATX headings as closing a block. - Keep a block open on an unended <!--, <?, <!X or <![CDATA[ line, on an open fence inside a list item or quote, and on a fence that has an earlier closing line marked skips. - Re-check a definition that is still streaming its line instead of rendering the rest of the message whole. - Move seededRandom to a dependency-free e2e/perf module so web tests no longer reach @thinkrail/server types. - List the remaining marked/micromark gaps in chat/SPEC.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ream event ChatView rebuilt recentPrompts from turns on every event, so Composer got a new array each time. Keep the previous array while its contents are equal. useChatTodos wrote a ref during render, so the React Compiler skipped it and it returned a new plan object every render, which re-rendered ChatHeader and the plan popover. Sync the ref in an effect and replace the try/catch value block with .catch(), so the hook compiles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Shiki's first highlight of each language ran on the main thread and blocked streaming for 200-450 ms at 4x CPU; on WebKit (the desktop webview) the JS regex engine costs ~1 s on a first TypeScript fence. - highlighter.worker.ts runs the chat highlighter off the main thread with the Oniguruma WASM engine; each grammar loads on first use. - highlightCode.ts is the client: per-key coalescing (one request in flight, latest call waits), a bounded LRU cache (256 entries / 4M chars, one entry per key, oversized results skipped), and a permanent in-thread fallback (JS regex engine) when no worker exists, it errors, or it cannot load the WASM engine. - useHighlightedCode() serves ShikiBlock and the tool CodeBlock: it reads the cache on the client only, so static renders (rendered diff) stay plain, applies replies in order, and shows the last highlight plus newer streamed text as plain until the worker catches up. - Vite emits workers as ES modules, so worker grammar imports split into on-demand chunks instead of inlining. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Andrey Markin (Mark-Life)
force-pushed
the
perf/shiki-worker
branch
from
October 5, 2026 15:23
b2bab48 to
7fd0332
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack (5/5): #616 profiling harness → #617 heavy-load scenarios → #618 streamed markdown blocks → #619 Composer re-renders → this PR.
Depends on #619. This PR targets
main, so it also shows the commits of #616–#619. Review only the last 2 commits:c8d4a520,5b92aa57.Problem
Streaming freezes for a moment each time the first code block of a new language arrives. With CPU throttled 4x, the first Python block blocks the main thread for about 250 ms, bash for about 180 ms, and TypeScript for 400–480 ms. In WebKit, which the desktop app runs on, the first TypeScript block freezes the page for 600–1,200 ms.
ShikiBlockcallshighlightCodeon the main thread. On first use of each language, Shiki's JavaScript regex engine converts and compiles that grammar's regexes in one synchronous task. Swapping the markdown renderer could not remove it: every renderer routes fences through the same call.Approach
Two changes. The bigger one is the engine: the chat highlighter now uses Shiki's Oniguruma WASM engine. On WebKit the first TypeScript highlight drops from about 880 ms to 47 ms, because JavaScriptCore is slow at running the converted regexes. On Chromium the WASM engine is about the same speed.
Then the work leaves the main thread. A module worker owns the highlighter; a small client sends keyed requests and the worker drops superseded ones. While a reply is pending, the block shows the last highlight plus the newer text as plain text, so streamed code never freezes on an old highlight. Results go into a size-bounded cache, so blocks remounted while scrolling history highlight instantly.
If the worker cannot start, for example a stale chunk after an update, the client switches to in-thread highlighting for good. Static renders (
renderToStaticMarkup, used by the rendered markdown diff) stay plain, as today.Changes
lib/highlighter.worker.ts,lib/highlightCode.ts: worker and client, withuseHighlightedCodeshared byShikiBlockand the toolCodeBlock.lib/highlighter.ts:createChatHighlighter(engine); the worker uses the WASM engine.vite.config.ts:worker.format: "es", so grammars and the WASM load on demand instead of being inlined. The Pierre worker shrinks from 301 KB to 69 KB gzip as a side effect. Monaco's worker is unchanged.lib,chatandapps/webSPECs.Cost: the worker graph carries its own copies of the grammars and the 231 KB gzip WASM. All of it loads lazily, only when code is highlighted.
Screenshots
Final markup is unchanged. Recording of the real app at the first TypeScript block, CPU 4x, played at 0.2x. The timer at the top is capture-only: before, it stops while Shiki blocks the main thread and the block then appears already colored; after, it keeps running and the code streams in, then turns colored.
shiki-worker-ts-recording.mp4
Checklist
bun run lint,bun run typecheck(17/17),bun run --cwd apps/web test(1,428 pass),check:deps,check:boundaries,check:seams,check:spec-surface.compiler:census: no new bail-outsbun run e2e): 452 passed, 1 failed.00-jbcentral.spec.ts:20fails the same way on the baseSPEC.md/ top-level specs updated to reflect any boundary, contract, or behavior changebun run perf:render --runs 1 --cpu 4, median of 5 interleaved runs, base vs this branch, Chromium:WebKit (Playwright, no throttle, 3 runs each), long-stream: longest frame gap 964 → 138 ms, gaps over 100 ms 3 → 1 per run.
Time from a fence's first text to colored code, py / bash / ts: Chromium at 1x 80 / 41 / 91 → 87 / 16 / 50 ms; WebKit 120 / 44 / 592 → 87 / 17 / 57 ms. Throttling does not slow workers, so the 1x row is the fair comparison.
Two numbers go the other way. React render time on long-stream rises 1448 → 1608 ms, because the unblocked thread commits more, smaller batches; total main-thread time is flat. The remaining ~140 ms long task is most likely mermaid's diagram layout (not profiled again here), and is not touched. Not run: the packaged Electrobun app. It serves the web build from
http://127.0.0.1, where the existing Pierre and htmldiff module workers already run.