Skip to content

@tanstack/ai-anthropic: a max_tokens stop emits RUN_ERROR with no token usage #1597

Description

@kyleromines

Package / version: @tanstack/ai-anthropic 0.19.3 (the same code is on main), @tanstack/ai latest.

Summary

When a streamed Anthropic call stops at max_tokens, the adapter emits RUN_ERROR with no usage. Anthropic reports usage for that call and bills it. The counts never reach the consumer.

Reproduction

import { chat } from '@tanstack/ai'
import { createAnthropicChat } from '@tanstack/ai-anthropic'

const adapter = createAnthropicChat('claude-haiku-4-5', process.env.ANTHROPIC_API_KEY)
for await (const c of chat({
  adapter,
  modelOptions: { max_tokens: 3 },
  messages: [{ role: 'user', content: 'Say hello in five words.' }],
})) {
  if (c.type === 'RUN_FINISHED' || c.type === 'RUN_ERROR') console.log(JSON.stringify(c))
}

Observed

{"type":"RUN_ERROR","message":"The response was cut off because the maximum token limit was reached.","code":"max_tokens","metadata":{"tanstack":{"model":"claude-haiku-4-5"}}}

There is no usage field. The same prompt with max_tokens: 20 ends in RUN_FINISHED with usage: { promptTokens: 13, completionTokens: 10, totalTokens: 23 }.

Anthropic sent usage for the truncated call. The raw stream's closing event was:

message_delta  stop_reason: max_tokens
usage: {"input_tokens":13,"cache_creation_input_tokens":0,"cache_read_input_tokens":0,"output_tokens":3}

The max_tokens branch in processAnthropicStream yields RUN_ERROR without calling buildAnthropicUsage(event.usage). The tool_use and default branches do call it.

Expected

The max_tokens RUN_ERROR carries the same usage as the other stop reasons. Alternatively, the usage is made available to the consumer some other way.

Why it matters

Anything that meters or bills on usage undercounts every call that ends this way. The call consumed input and output tokens, but the consumer has no count for it.

Secondary note

buildAnthropicUsage reads a missing input_tokens on the closing message_delta as 0 (usage.input_tokens ?? 0). Anthropic's API currently repeats all input counts on the closing delta, so I could not reproduce a problem against it. Some Anthropic-compatible servers may send only output_tokens there. For those, input and cache counts would read as 0, where Anthropic's own SDK keeps the message_start values. I haven't tested this, so treat it as speculative. A merge of message_start and message_delta usage would cover it.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

has-prAn open PR references this issuewaiting-on: maintainerThe ball is in the maintainers’ court

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions