Skip to content

feat(tokens): account for image inputs in token usage - #75

Merged
chaliy merged 1 commit into
mainfrom
claude/llmsim-multimodal-inputs-8gpnha
Jun 28, 2026
Merged

chaliy merged 1 commit into
mainfrom
claude/llmsim-multimodal-inputs-8gpnha

Conversation

@chaliy

@chaliy chaliy commented Jun 28, 2026

Copy link
Copy Markdown
Owner

What

Image inputs now contribute an approximate token cost to reported usage, across every endpoint that accepts images:

  • Chat Completions → prompt_tokens
  • OpenAI Responses API → input_tokens
  • OpenResponses → input_tokens

Follow-up to #73 (which made the endpoints accept and gate multimodal input). This makes usage reflect images instead of silently ignoring them.

Why

After #73, image parts were accepted and validated but dropped before token counting, so a request with images reported the same token usage as a text-only request. Clients that budget against usage (cost estimation, rate-limit simulation) saw no signal from image content.

How

  • New estimate_image_tokens(detail) in src/tokens.rs (+ IMAGE_TOKENS_LOW/IMAGE_TOKENS_HIGH constants), exported from the crate root.
  • Decision: the simulator never fetches or decodes image bytes, so the real tile-based count (which needs pixel dimensions) is impossible. We approximate from the detail hint using OpenAI's documented gpt-4o costs: "low" → 85 tokens; "high"/"auto"/unset → 765 (a representative ~1024×1024 image, 85 + 4*170).
  • Chat Completions count_request_tokens adds per-image cost from each part's detail.
  • count_responses_input_image_tokens (Responses) and count_openresponses_input_image_tokens (OpenResponses) fold image cost into input-token counts. The OpenAI Responses input_image part has no detail, so it uses the high-detail default.
  • Image content still does not influence the generated output text.
  • Specs (api-endpoints.md, responses-api.md) and docs (docs/api.md) updated.

Risk

  • Low
  • Only affects reported token counts when image parts are present; text-only requests are unchanged. Approximation is documented as such. Verified end-to-end: high-detail image adds 765, low-detail adds 85 across all three endpoints.

Checklist

  • Tests added or updated (estimate_image_tokens; per-endpoint image-token accounting)
  • Backward compatibility considered (text-only usage unchanged)

Generated by Claude Code

Estimate per-image token cost and fold it into reported usage across all
endpoints that accept images: chat completions (prompt_tokens), the OpenAI
Responses API and OpenResponses (input_tokens).

The simulator never fetches or decodes image bytes, so the real tile-based
count is impossible. estimate_image_tokens approximates from the detail hint
using OpenAI's documented gpt-4o costs: low -> 85 tokens, high/auto/unset ->
765 (a representative ~1024x1024 image). Responses input_image parts carry
no detail and use the high default. Image content still does not influence
generated output text.
@chaliy
chaliy merged commit 14cf904 into main Jun 28, 2026
11 checks passed
@chaliy
chaliy deleted the claude/llmsim-multimodal-inputs-8gpnha branch June 28, 2026 16:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant