Trasys
Node.js SDK

Health Checks & Alerts

Register named health check functions on a schedule, define alert rule severities, and trigger autonomous incident response when a check starts failing.

Health Checks & Alerts

Health checks

Health checks are named functions that run on a schedule. Results are sent to the Trasys platform and can trigger alerts when a check starts failing.

Registering a check

const { sdk } = require('./trasys');

sdk.healthCheck('database', async () => {
  await pool.query('SELECT 1');
}, {
  timeoutMs:  3000,  // fail the check if it takes longer than 3s
  intervalMs: 30000, // run every 30 seconds
  severity:   'P1',  // alert severity when this check fails
});

The check function should:

  • Return (resolve) to signal healthy
  • Throw to signal unhealthy — the thrown error message is captured

Chain multiple checks with method chaining:

sdk
  .healthCheck('database', () => pool.query('SELECT 1'))
  .healthCheck('redis',    () => redis.ping())
  .healthCheck('stripe',   () => stripe.balance.retrieve());

Options

OptionTypeDefaultDescription
timeoutMsnumber5000Milliseconds before the check is considered timed out
intervalMsnumber30000How often to run the check
severity'P1' | 'P2' | 'P3''P2'Alert severity when the check fails

Exposing a health endpoint

app.get('/health', async (req, res) => {
  const status = await sdk.getHealthStatus();
  res.status(status.status === 'healthy' ? 200 : 503).json(status);
});

Response shape:

{
  "status": "degraded",
  "checks": [
    { "name": "database", "status": "healthy",   "latencyMs": 4   },
    { "name": "redis",    "status": "unhealthy",  "latencyMs": 0,  "error": "ECONNREFUSED" },
    { "name": "stripe",   "status": "healthy",   "latencyMs": 120 }
  ]
}

Overall status logic:

StatusMeaning
healthyAll checks passed
degradedSome passed, some failed
unhealthyAll checks failed

Alerts

Alert rules define conditions that, when met, send notifications and optionally trigger the Trasys SRE agent.

Defining a rule

sdk.defineAlert({
  id:          'high-error-rate',
  name:        'High Error Rate',
  description: 'Error rate exceeded 5% for 2 consecutive minutes',
  condition:   'error_rate("payment-service") > 0.05',

  windowSeconds:             120, // evaluate over a 2-minute window
  evaluationIntervalSeconds: 30,  // check every 30 seconds
  severity:                  'P1',

  notify: {
    channels: ['slack:#incidents', 'pagerduty'],
    escalation: [
      { afterMinutes: 5, to: 'on-call-engineer', via: ['sms', 'phone'] },
    ],
  },

  environment:     'production',
  cooldownSeconds: 900, // don't re-fire for 15 minutes after resolving
  runbook:         'rb-payment-errors',
});

Alert severity

SeverityNotificationSRE agent
P1SMS + phone call immediatelyClaude Opus (extended thinking)
P2Slack immediately; SMS after 5 minClaude Sonnet
P3Slack onlyClaude Haiku

See AI SRE Agent for exactly what the agent does with a firing alert, the full alert condition function reference, and how rules sync from your code to the backend.

Defining multiple rules

sdk.defineAlerts([
  {
    id:        'p99-latency',
    name:      'P99 Latency Spike',
    condition: 'p99_latency("api-gateway") > 2000',
    severity:  'P2',
    notify:    { channels: ['slack:#performance'] },
  },
  {
    id:        'ai-cost-overrun',
    name:      'AI Daily Cost Limit',
    condition: 'daily_cost() > 500',
    severity:  'P3',
    notify:    { channels: ['slack:#engineering'] },
  },
]);

Alert condition reference

Service health

error_rate("service-name")    > 0.05    // error rate > 5%
success_rate("service-name")  < 0.95    // success rate < 95%
p50_latency("service-name")   > 500     // median latency > 500ms
p95_latency("service-name")   > 2000    // p95 latency > 2s
p99_latency("service-name")   > 5000    // p99 latency > 5s
request_rate("service-name")  < 10      // fewer than 10 req/s (drop detection)

AI and agents

daily_cost()                       > 500     // total AI cost today > $500
daily_cost("agent-name")           > 100     // cost for a specific agent
token_cost_per_run("agent-name")   > 0.50    // cost per agent run > $0.50
tool_failure_rate("agent", "tool") > 0.10    // tool failure rate > 10%
agent_loop_detected("agent-name")            // agent entered an infinite loop

Health checks

health_check_status("database") == "unhealthy"

Combining conditions

p99_latency("checkout") > 3000 AND error_rate("checkout") > 0.02

Notification channels

FormatExample
slack:#channel'slack:#incidents'
slack:@user'slack:@on-call'
email:address'email:alerts@company.com'
sms'sms'
phone'phone'
pagerduty'pagerduty'

Escalation

notify: {
  channels: ['slack:#incidents'],
  escalation: [
    { afterMinutes: 0,  to: 'primary-oncall',  via: ['sms'] },
    { afterMinutes: 10, to: 'secondary-oncall', via: ['sms', 'phone'] },
    { afterMinutes: 20, to: 'engineering-lead', via: ['phone'] },
  ],
},

afterMinutes: 0 fires at the same time as the initial notification. Higher values escalate if the alert is not acknowledged within that time.


Next steps

  • AI SRE Agent — what happens after an alert fires, and the full condition function reference
  • TQL — query health check results over time with FROM health_checks
  • Session Tracking — correlate health failures with specific user activity

On this page