feat(tokens): account for image inputs in token usage - #75
Merged
Merged
Conversation
Estimate per-image token cost and fold it into reported usage across all endpoints that accept images: chat completions (prompt_tokens), the OpenAI Responses API and OpenResponses (input_tokens). The simulator never fetches or decodes image bytes, so the real tile-based count is impossible. estimate_image_tokens approximates from the detail hint using OpenAI's documented gpt-4o costs: low -> 85 tokens, high/auto/unset -> 765 (a representative ~1024x1024 image). Responses input_image parts carry no detail and use the high default. Image content still does not influence generated output text.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Image inputs now contribute an approximate token cost to reported
usage, across every endpoint that accepts images:prompt_tokensinput_tokensinput_tokensFollow-up to #73 (which made the endpoints accept and gate multimodal input). This makes
usagereflect images instead of silently ignoring them.Why
After #73, image parts were accepted and validated but dropped before token counting, so a request with images reported the same token usage as a text-only request. Clients that budget against
usage(cost estimation, rate-limit simulation) saw no signal from image content.How
estimate_image_tokens(detail)insrc/tokens.rs(+IMAGE_TOKENS_LOW/IMAGE_TOKENS_HIGHconstants), exported from the crate root.detailhint using OpenAI's documented gpt-4o costs:"low"→ 85 tokens;"high"/"auto"/unset → 765 (a representative ~1024×1024 image,85 + 4*170).count_request_tokensadds per-image cost from each part'sdetail.count_responses_input_image_tokens(Responses) andcount_openresponses_input_image_tokens(OpenResponses) fold image cost into input-token counts. The OpenAI Responsesinput_imagepart has nodetail, so it uses the high-detail default.api-endpoints.md,responses-api.md) and docs (docs/api.md) updated.Risk
Checklist
estimate_image_tokens; per-endpoint image-token accounting)Generated by Claude Code