Trasys
Node.js SDK

AI SRE Agent

How Trasys's autonomous incident response works — alert severity maps to a specific Claude model, the SDK defines rules, and the backend investigates and triggers the agent.

AI SRE Agent

Every alert rule you define with sdk.defineAlert() doesn't just notify a human — it can also trigger Trasys's AI SRE agent to investigate the incident automatically. The model that responds scales with how severe the alert is.

This page covers what the SRE agent is and how severity maps to it. For the mechanics of defining alert rules themselves — conditions, notification channels, escalation — see Health Checks & Alerts.

What the SDK does vs. what the backend does

This split matters for understanding what actually happens when an alert fires:

  1. Your code calls sdk.defineAlert(rule) — the rule is validated and held in memory.
  2. The SDK syncs every defined rule to the backend (POST /v1/alert-rules) at startup and again every 5 minutes, so rules stay in sync across deploys with no dashboard step required.
  3. The backend — not the SDK — evaluates every rule every 30 seconds against your ClickHouse data.
  4. On a rule firing, the backend creates an incident and triggers the SRE agent.

The SDK never evaluates alert conditions locally and never talks to the SRE agent directly — its only job is defining and syncing the rule.

Severity determines the model

Each AlertSeverity ('P1' | 'P2' | 'P3') maps to both a notification urgency and a specific Claude model doing the investigation:

SeverityMeaningNotificationSRE agent model
P1Critical — production down or severely degradedImmediate SMS + phone callclaude-opus-4-6, with extended thinking
P2High — significant impact, not a total outageSlack immediately; SMS after 5 minclaude-sonnet-4-6
P3Medium — minor impact or an early warning signalSlack onlyclaude-haiku-4-5

The idea: a full outage gets your most capable model doing deep investigation, while a minor early-warning signal gets a fast, cheap triage pass. You don't choose the model directly — you choose the severity, and the agent tier follows from it.

Defining a rule

sdk.defineAlert({
  id:          'payment-high-error-rate',
  name:        'Payment Service High Error Rate',
  description: 'Error rate exceeded 5% for 2 consecutive minutes',
  condition:   "error_rate('payment-service') > 0.05",

  windowSeconds:             120, // condition must hold for 2 minutes before firing
  evaluationIntervalSeconds: 30,  // backend checks every 30s
  severity:                  'P1',

  notify: {
    channels: ['slack:#incidents', 'pagerduty'],
    escalation: [
      { afterMinutes: 0,  to: 'primary-oncall',   via: ['sms'] },
      { afterMinutes: 5,  to: 'secondary-oncall', via: ['sms', 'phone'] },
    ],
  },

  runbook: 'payment-errors-runbook', // agent starts here instead of investigating from scratch
});

runbook is optional — if you don't provide one, the agent investigates from scratch using whatever telemetry it has access to.

Alert condition functions

condition is a string evaluated server-side against ClickHouse. Every function below is built in:

error_rate(service, filters?)       → decimal 0.0–1.0
success_rate(service, filters?)     → decimal 0.0–1.0
p50_latency(service)                → milliseconds
p95_latency(service)                → milliseconds
p99_latency(service)                → milliseconds
request_rate(service)               → requests per second
daily_cost(agent?)                  → USD, account-wide or scoped to one AI agent
token_cost_per_run(agent)           → USD per agent run
tool_failure_rate(agent, tool?)     → decimal 0.0–1.0
hallucination_rate(agent)           → decimal 0.0–1.0 (beta)
agent_loop_detected(agent)          → boolean (0 or 1)
health_check_status(name)           → 0 = healthy, 1 = unhealthy

The agent_loop_detected() and tool_failure_rate() functions are specifically for monitoring AI agents that make tool calls — separate from error_rate/p95_latency, which apply to any HTTP service. Combine conditions with AND/OR:

"p99_latency('checkout') > 3000 AND error_rate('checkout') > 0.02"

Extra context for the agent

Pass anything that would help a human (or the agent) triage faster — dashboard links, the right Slack channel, the on-call rotation name:

sdk.defineAlert({
  id: 'invoice-agent-cost-spike',
  // ...
  context: {
    dashboardLink:  'https://app.trasys.dev/agents/invoice-processor',
    slackChannel:   '#payments-team',
    oncallRotation: 'payments-oncall',
  },
});

Deduplication and updates

id is the dedup key. Calling defineAlert() again with the same id — on every deploy, since your alert rules live in your source code — updates the existing rule rather than creating a duplicate. There's no dashboard-side rule authoring required; rules are versioned alongside your application code.

Cooldown

By default, a resolved alert can't re-fire for 15 minutes (cooldownSeconds: 900), so a condition that oscillates around its threshold doesn't spam notifications or re-trigger the agent repeatedly for what's really one ongoing issue.


Next steps

  • Health Checks & Alerts — health check registration, the full notification channel list, and escalation policies
  • TQL — the same aggregation functions available in alert conditions (error_rate, p95_latency, etc.) are queryable directly
  • Metrics — custom monitor.* metrics that feed your own dashboards alongside the built-in alert functions

On this page