Tests that call a real model are slow, cost money, need an API key, and get a different answer each time. fakeText() from @tanstack/ai/testing fixes that. It is a text adapter that answers from a script you write. Pass it to chat() like any other adapter.
Queue the answers, then run chat():
import { chat } from '@tanstack/ai'
import { fakeText } from '@tanstack/ai/testing'
const fake = fakeText()
fake.setResponses([{ text: 'Hello!' }])
let answer = ''
for await (const chunk of chat({
adapter: fake,
messages: [{ role: 'user', content: 'Hi' }],
})) {
if (chunk.type === 'TEXT_MESSAGE_CONTENT') answer += chunk.delta
}
console.log(answer) // 'Hello!'Give an answer with toolCalls, and chat() runs your tool. Then queue the answer that comes after the tool result:
import { toolDefinition } from '@tanstack/ai'
import { z } from 'zod'
const getWeather = toolDefinition({
name: 'get_weather',
description: 'Get the weather in a city',
inputSchema: z.object({ city: z.string() }),
}).server(async ({ city }) => ({ city, sky: 'sunny' }))
const weatherFake = fakeText()
weatherFake.setResponses([
{ toolCalls: [{ name: 'get_weather', input: { city: 'Oslo' } }] },
{ text: 'It is sunny in Oslo.' },
])
for await (const chunk of chat({
adapter: weatherFake,
messages: [{ role: 'user', content: 'Weather in Oslo?' }],
tools: [getWeather],
})) {
if (chunk.type === 'TOOL_CALL_RESULT') console.log(chunk.content)
}
console.log(weatherFake.state.callCount) // 2Put a function in the queue to build the answer when the call comes in. It gets the request and the state:
const echo = fakeText()
echo.setResponses([
({ request, state }) => ({
text: `Call ${state.callCount} saw ${request.messages.length} messages.`,
}),
])Give an answer with error, and the call fails with a RUN_ERROR:
const broken = fakeText()
broken.setResponses([{ error: 'The provider is down' }])
for await (const chunk of chat({
adapter: broken,
messages: [{ role: 'user', content: 'Hi' }],
})) {
if (chunk.type === 'RUN_ERROR') console.log(chunk.message) // 'The provider is down'
}The fake also reports token usage on each call: ceil(characters / 4) over the request and the answer. So a long message costs many tokens, like with a real model.
| Option | What it does |
|---|---|
| model | The model id. Default 'fake-model'. |
| contextWindow | The context window in tokens. Read it back as fake.contextWindow. |
| tokensPerSecond | Stream the text at this speed, 4 characters per token. |
| cache | With a threadId, the part of the request that matches the previous request of the thread counts as cached tokens. |
| Field | What it does |
|---|---|
| text | The visible answer. |
| thinking | Thinking text, streamed before the answer. |
| toolCalls | { name, input?, id? } calls for chat() to run. |
| finishReason | 'stop', 'length', 'content_filter', or 'tool_calls'. Default 'tool_calls' with tool calls, else 'stop'. |
| error | Fail the call with a RUN_ERROR that has this message. |
Your tests now run with no network and no key, and they get the same answer every time.