Trasys
Node.js SDK

AI Observability

Trace every LLM call, model, tokens, cost, prompts, responses, tool calls, across OpenAI, Anthropic, Gemini, Groq, Mistral, Cohere, and Ollama via sdk.wrapAI.

AI Observability

The SDK instruments AI provider clients to capture every LLM call as a span — model name, token usage, estimated cost, prompts, responses, tool calls, and finish reason. Streaming responses are fully supported.

Wrapping a client

All providers use the same API: pass your client instance to sdk.wrapAI(). It returns the same client with instrumentation applied — no type changes, no interface changes.

const { sdk } = require('./trasys');
const OpenAI  = require('openai');

const openai = sdk.wrapAI(new OpenAI({ apiKey: process.env.OPENAI_API_KEY }));

// Use exactly as before — all calls are now traced
const completion = await openai.chat.completions.create({ ... });

sdk.wrapAI() detects the provider automatically by inspecting the client instance (e.g. client.chat.completions.create means OpenAI-compatible, client.messages.create without .chat means Anthropic). If detection fails — an unrecognized or very new client shape — it logs a warning and returns the client unwrapped rather than guessing wrong. Use a provider-specific wrapper directly if you want to skip detection entirely:

const { wrapOpenAI, wrapAnthropic, wrapGemini } = require('@trasys/sdk');

const openai    = wrapOpenAI(new OpenAI());
const anthropic = wrapAnthropic(new Anthropic());

Per-client options

const openai = sdk.wrapAI(new OpenAI(), {
  agentName:     'support-bot', // label shown in dashboard; defaults to service name
  captureInput:  true,          // override ai.capturePrompts for this client
  captureOutput: true,          // override ai.captureResponses for this client
});

OpenAI

const { sdk } = require('./trasys');
const OpenAI  = require('openai');

const openai = sdk.wrapAI(new OpenAI());

// Non-streaming
const response = await openai.chat.completions.create({
  model:    'gpt-4o',
  messages: [{ role: 'user', content: 'Summarize this article...' }],
});

// Streaming
const stream = await openai.chat.completions.create({
  model:    'gpt-4o',
  messages: [{ role: 'user', content: 'Tell me a story.' }],
  stream:   true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? '');
}

For streaming, the SDK automatically adds stream_options: { include_usage: true } so OpenAI returns token counts in the final chunk. The span is finalized when the stream closes.


Anthropic

const { sdk }   = require('./trasys');
const Anthropic = require('@anthropic-ai/sdk');

const anthropic = sdk.wrapAI(new Anthropic());

// Non-streaming
const message = await anthropic.messages.create({
  model:      'claude-sonnet-4-6',
  max_tokens: 1024,
  messages:   [{ role: 'user', content: 'Explain async/await.' }],
});

// Streaming
const stream = await anthropic.messages.create({
  model:      'claude-sonnet-4-6',
  max_tokens: 1024,
  messages:   [{ role: 'user', content: 'Write a poem.' }],
  stream:     true,
});

for await (const event of stream) {
  if (event.type === 'content_block_delta') {
    process.stdout.write(event.delta.text);
  }
}

For streaming, the SDK accumulates token counts from message_start (input) and message_delta (output + stop reason). Anthropic prompt cache tokens (cache_creation_input_tokens, cache_read_input_tokens) are captured when present.


Google Gemini

const { sdk }                = require('./trasys');
const { GoogleGenerativeAI } = require('@google/generative-ai');

const genAI = sdk.wrapAI(new GoogleGenerativeAI(process.env.GEMINI_API_KEY));
const model = genAI.getGenerativeModel({ model: 'gemini-1.5-pro' });

// Non-streaming
const result = await model.generateContent('What is the capital of France?');
console.log(result.response.text());

// Streaming
const streamResult = await model.generateContentStream('Write a haiku.');
for await (const chunk of streamResult.stream) {
  process.stdout.write(chunk.text());
}

// Chat / RAG implementations
const chat = model.startChat({ history: [] });
const msg  = await chat.sendMessage('Hello!');

If you use model.startChat() for RAG pipelines, all sendMessage() and sendMessageStream() calls are traced automatically. No changes to your RAG code are needed.


Groq

const { sdk } = require('./trasys');
const Groq    = require('groq-sdk');

const groq = sdk.wrapAI(new Groq({ apiKey: process.env.GROQ_API_KEY }));

const completion = await groq.chat.completions.create({
  model:    'llama3-8b-8192',
  messages: [{ role: 'user', content: 'Hello!' }],
});

Groq's API mirrors OpenAI's interface and the instrumentation behaves identically.


Mistral

const { sdk }     = require('./trasys');
const { Mistral } = require('@mistralai/mistralai');

const mistral = sdk.wrapAI(new Mistral({ apiKey: process.env.MISTRAL_API_KEY }));

// Non-streaming
const result = await mistral.chat.complete({
  model:    'mistral-large-latest',
  messages: [{ role: 'user', content: 'Explain recursion.' }],
});

// Streaming
const stream = await mistral.chat.stream({
  model:    'mistral-large-latest',
  messages: [{ role: 'user', content: 'Write a summary.' }],
});

for await (const chunk of stream) {
  process.stdout.write(chunk.data.choices[0]?.delta?.content ?? '');
}

Cohere

const { sdk }          = require('./trasys');
const { CohereClient } = require('cohere-ai');

const cohere = sdk.wrapAI(new CohereClient({ token: process.env.COHERE_API_KEY }));

// Non-streaming
const response = await cohere.chat({
  model:   'command-r-plus',
  message: 'Summarize this document...',
});

// Streaming
const stream = await cohere.chatStream({
  model:   'command-r-plus',
  message: 'Explain machine learning.',
});

for await (const event of stream) {
  if (event.eventType === 'text-generation') {
    process.stdout.write(event.text);
  }
}

For streaming, usage metadata is read from the stream-end event which carries the full response summary.


Ollama

const { sdk }    = require('./trasys');
const { Ollama } = require('ollama');

const ollama = sdk.wrapAI(new Ollama({ host: 'http://localhost:11434' }));

const response = await ollama.chat({
  model:    'llama3',
  messages: [{ role: 'user', content: 'Hello!' }],
});

What gets captured per call

AttributeDescription
gen_ai.systemProvider — openai, anthropic, gemini…
gen_ai.request.modelModel name
gen_ai.usage.input_tokensPrompt token count
gen_ai.usage.output_tokensCompletion token count
gen_ai.response.finish_reasonsstop, length, tool_calls…
gen_ai.tool_callsTool/function calls made during the response
trasys.ai.cost_usdEstimated cost in USD
trasys.ai.promptEncrypted prompt (when capturePrompts: true)
trasys.ai.responseEncrypted response (when captureResponses: true)

Prompt and response capture

Prompts and responses are captured and encrypted with AES-256-GCM before transmission. The encryption key is derived from your API key.

See Prompt & Response Data Security for the full capture/masking config reference and the encryption utility's API.

To disable capture globally:

createSdk({
  ai: {
    capturePrompts:   false,
    captureResponses: false,
  },
});

To disable per client:

const openai = sdk.wrapAI(new OpenAI(), {
  captureInput:  false,
  captureOutput: true,
});

To mask specific fields before encryption:

createSdk({
  ai: { maskFields: ['password', 'ssn', 'credit_card'] },
});

Next steps

  • Prompt & Response Data Security — capture toggles, field masking, and the encryption utility
  • TQL — query AI spans by cost, model, token usage, and more
  • AI SRE Agent — agent_loop_detected() and tool_failure_rate() alert conditions built specifically for AI agents

On this page