TanStack
Advanced

Test with a Fake Model

Tests that call a real model are slow, cost money, need an API key, and get a different answer each time. fakeText() from @tanstack/ai/testing fixes that. It is a text adapter that answers from a script you write. Pass it to chat() like any other adapter.

Answer a prompt

Queue the answers, then run chat():

ts
import { chat } from '@tanstack/ai'
import { fakeText } from '@tanstack/ai/testing'

const fake = fakeText()
fake.setResponses([{ text: 'Hello!' }])

let answer = ''
for await (const chunk of chat({
  adapter: fake,
  messages: [{ role: 'user', content: 'Hi' }],
})) {
  if (chunk.type === 'TEXT_MESSAGE_CONTENT') answer += chunk.delta
}
console.log(answer) // 'Hello!'
  • Each model call takes the next answer from the queue.
  • An empty queue ends the call with a RUN_ERROR: "No more fake responses queued".
  • appendResponses adds answers. pendingResponses() counts what is left.

Test a tool

Give an answer with toolCalls, and chat() runs your tool. Then queue the answer that comes after the tool result:

ts
import { toolDefinition } from '@tanstack/ai'
import { z } from 'zod'

const getWeather = toolDefinition({
  name: 'get_weather',
  description: 'Get the weather in a city',
  inputSchema: z.object({ city: z.string() }),
}).server(async ({ city }) => ({ city, sky: 'sunny' }))

const weatherFake = fakeText()
weatherFake.setResponses([
  { toolCalls: [{ name: 'get_weather', input: { city: 'Oslo' } }] },
  { text: 'It is sunny in Oslo.' },
])

for await (const chunk of chat({
  adapter: weatherFake,
  messages: [{ role: 'user', content: 'Weather in Oslo?' }],
  tools: [getWeather],
})) {
  if (chunk.type === 'TOOL_CALL_RESULT') console.log(chunk.content)
}
console.log(weatherFake.state.callCount) // 2

Answer from the request

Put a function in the queue to build the answer when the call comes in. It gets the request and the state:

ts
const echo = fakeText()
echo.setResponses([
  ({ request, state }) => ({
    text: `Call ${state.callCount} saw ${request.messages.length} messages.`,
  }),
])

Test an error

Give an answer with error, and the call fails with a RUN_ERROR:

ts
const broken = fakeText()
broken.setResponses([{ error: 'The provider is down' }])

for await (const chunk of chat({
  adapter: broken,
  messages: [{ role: 'user', content: 'Hi' }],
})) {
  if (chunk.type === 'RUN_ERROR') console.log(chunk.message) // 'The provider is down'
}

The fake also reports token usage on each call: ceil(characters / 4) over the request and the answer. So a long message costs many tokens, like with a real model.

Options

OptionWhat it does
modelThe model id. Default 'fake-model'.
contextWindowThe context window in tokens. Read it back as fake.contextWindow.
tokensPerSecondStream the text at this speed, 4 characters per token.
cacheWith a threadId, the part of the request that matches the previous request of the thread counts as cached tokens.

Answer fields

FieldWhat it does
textThe visible answer.
thinkingThinking text, streamed before the answer.
toolCalls{ name, input?, id? } calls for chat() to run.
finishReason'stop', 'length', 'content_filter', or 'tool_calls'. Default 'tool_calls' with tool calls, else 'stop'.
errorFail the call with a RUN_ERROR that has this message.

Your tests now run with no network and no key, and they get the same answer every time.