AI Observability
Trace every LLM call, model, tokens, cost, prompts, responses, tool calls, across OpenAI, Anthropic, Gemini, Groq, Mistral, Cohere, and Ollama via sdk.wrapAI.
AI Observability
The SDK instruments AI provider clients to capture every LLM call as a span — model name, token usage, estimated cost, prompts, responses, tool calls, and finish reason. Streaming responses are fully supported.
Wrapping a client
All providers use the same API: pass your client instance to sdk.wrapAI(). It returns the same client with instrumentation applied — no type changes, no interface changes.
const { sdk } = require('./trasys');
const OpenAI = require('openai');
const openai = sdk.wrapAI(new OpenAI({ apiKey: process.env.OPENAI_API_KEY }));
// Use exactly as before — all calls are now traced
const completion = await openai.chat.completions.create({ ... });sdk.wrapAI() detects the provider automatically by inspecting the client instance (e.g. client.chat.completions.create means OpenAI-compatible, client.messages.create without .chat means Anthropic). If detection fails — an unrecognized or very new client shape — it logs a warning and returns the client unwrapped rather than guessing wrong. Use a provider-specific wrapper directly if you want to skip detection entirely:
const { wrapOpenAI, wrapAnthropic, wrapGemini } = require('@trasys/sdk');
const openai = wrapOpenAI(new OpenAI());
const anthropic = wrapAnthropic(new Anthropic());Per-client options
const openai = sdk.wrapAI(new OpenAI(), {
agentName: 'support-bot', // label shown in dashboard; defaults to service name
captureInput: true, // override ai.capturePrompts for this client
captureOutput: true, // override ai.captureResponses for this client
});OpenAI
const { sdk } = require('./trasys');
const OpenAI = require('openai');
const openai = sdk.wrapAI(new OpenAI());
// Non-streaming
const response = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [{ role: 'user', content: 'Summarize this article...' }],
});
// Streaming
const stream = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [{ role: 'user', content: 'Tell me a story.' }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? '');
}For streaming, the SDK automatically adds stream_options: { include_usage: true } so OpenAI returns token counts in the final chunk. The span is finalized when the stream closes.
Anthropic
const { sdk } = require('./trasys');
const Anthropic = require('@anthropic-ai/sdk');
const anthropic = sdk.wrapAI(new Anthropic());
// Non-streaming
const message = await anthropic.messages.create({
model: 'claude-sonnet-4-6',
max_tokens: 1024,
messages: [{ role: 'user', content: 'Explain async/await.' }],
});
// Streaming
const stream = await anthropic.messages.create({
model: 'claude-sonnet-4-6',
max_tokens: 1024,
messages: [{ role: 'user', content: 'Write a poem.' }],
stream: true,
});
for await (const event of stream) {
if (event.type === 'content_block_delta') {
process.stdout.write(event.delta.text);
}
}For streaming, the SDK accumulates token counts from message_start (input) and message_delta (output + stop reason). Anthropic prompt cache tokens (cache_creation_input_tokens, cache_read_input_tokens) are captured when present.
Google Gemini
const { sdk } = require('./trasys');
const { GoogleGenerativeAI } = require('@google/generative-ai');
const genAI = sdk.wrapAI(new GoogleGenerativeAI(process.env.GEMINI_API_KEY));
const model = genAI.getGenerativeModel({ model: 'gemini-1.5-pro' });
// Non-streaming
const result = await model.generateContent('What is the capital of France?');
console.log(result.response.text());
// Streaming
const streamResult = await model.generateContentStream('Write a haiku.');
for await (const chunk of streamResult.stream) {
process.stdout.write(chunk.text());
}
// Chat / RAG implementations
const chat = model.startChat({ history: [] });
const msg = await chat.sendMessage('Hello!');If you use model.startChat() for RAG pipelines, all sendMessage() and sendMessageStream() calls are traced automatically. No changes to your RAG code are needed.
Groq
const { sdk } = require('./trasys');
const Groq = require('groq-sdk');
const groq = sdk.wrapAI(new Groq({ apiKey: process.env.GROQ_API_KEY }));
const completion = await groq.chat.completions.create({
model: 'llama3-8b-8192',
messages: [{ role: 'user', content: 'Hello!' }],
});Groq's API mirrors OpenAI's interface and the instrumentation behaves identically.
Mistral
const { sdk } = require('./trasys');
const { Mistral } = require('@mistralai/mistralai');
const mistral = sdk.wrapAI(new Mistral({ apiKey: process.env.MISTRAL_API_KEY }));
// Non-streaming
const result = await mistral.chat.complete({
model: 'mistral-large-latest',
messages: [{ role: 'user', content: 'Explain recursion.' }],
});
// Streaming
const stream = await mistral.chat.stream({
model: 'mistral-large-latest',
messages: [{ role: 'user', content: 'Write a summary.' }],
});
for await (const chunk of stream) {
process.stdout.write(chunk.data.choices[0]?.delta?.content ?? '');
}Cohere
const { sdk } = require('./trasys');
const { CohereClient } = require('cohere-ai');
const cohere = sdk.wrapAI(new CohereClient({ token: process.env.COHERE_API_KEY }));
// Non-streaming
const response = await cohere.chat({
model: 'command-r-plus',
message: 'Summarize this document...',
});
// Streaming
const stream = await cohere.chatStream({
model: 'command-r-plus',
message: 'Explain machine learning.',
});
for await (const event of stream) {
if (event.eventType === 'text-generation') {
process.stdout.write(event.text);
}
}For streaming, usage metadata is read from the stream-end event which carries the full response summary.
Ollama
const { sdk } = require('./trasys');
const { Ollama } = require('ollama');
const ollama = sdk.wrapAI(new Ollama({ host: 'http://localhost:11434' }));
const response = await ollama.chat({
model: 'llama3',
messages: [{ role: 'user', content: 'Hello!' }],
});What gets captured per call
| Attribute | Description |
|---|---|
gen_ai.system | Provider — openai, anthropic, gemini… |
gen_ai.request.model | Model name |
gen_ai.usage.input_tokens | Prompt token count |
gen_ai.usage.output_tokens | Completion token count |
gen_ai.response.finish_reasons | stop, length, tool_calls… |
gen_ai.tool_calls | Tool/function calls made during the response |
trasys.ai.cost_usd | Estimated cost in USD |
trasys.ai.prompt | Encrypted prompt (when capturePrompts: true) |
trasys.ai.response | Encrypted response (when captureResponses: true) |
Prompt and response capture
Prompts and responses are captured and encrypted with AES-256-GCM before transmission. The encryption key is derived from your API key.
See Prompt & Response Data Security for the full capture/masking config reference and the encryption utility's API.
To disable capture globally:
createSdk({
ai: {
capturePrompts: false,
captureResponses: false,
},
});To disable per client:
const openai = sdk.wrapAI(new OpenAI(), {
captureInput: false,
captureOutput: true,
});To mask specific fields before encryption:
createSdk({
ai: { maskFields: ['password', 'ssn', 'credit_card'] },
});Next steps
- Prompt & Response Data Security — capture toggles, field masking, and the encryption utility
- TQL — query AI spans by cost, model, token usage, and more
- AI SRE Agent —
agent_loop_detected()andtool_failure_rate()alert conditions built specifically for AI agents
Database Instrumentation
Instrument pg, Prisma, Drizzle, Mongoose, MongoDB, and Sequelize with the Trasys SDK — automatic for most clients, with manual wrappers for Prisma and Drizzle.
Prompt & Response Data Security
Control what AI prompt and response content the SDK captures, mask sensitive fields before it's recorded, and use the SDK's AES-256-GCM encryption utility for your own captured payloads.

