Health Checks & Alerts
Register named health check functions on a schedule, define alert rule severities, and trigger autonomous incident response when a check starts failing.
Health Checks & Alerts
Health checks
Health checks are named functions that run on a schedule. Results are sent to the Trasys platform and can trigger alerts when a check starts failing.
Registering a check
const { sdk } = require('./trasys');
sdk.healthCheck('database', async () => {
await pool.query('SELECT 1');
}, {
timeoutMs: 3000, // fail the check if it takes longer than 3s
intervalMs: 30000, // run every 30 seconds
severity: 'P1', // alert severity when this check fails
});The check function should:
- Return (resolve) to signal healthy
- Throw to signal unhealthy — the thrown error message is captured
Chain multiple checks with method chaining:
sdk
.healthCheck('database', () => pool.query('SELECT 1'))
.healthCheck('redis', () => redis.ping())
.healthCheck('stripe', () => stripe.balance.retrieve());Options
| Option | Type | Default | Description |
|---|---|---|---|
timeoutMs | number | 5000 | Milliseconds before the check is considered timed out |
intervalMs | number | 30000 | How often to run the check |
severity | 'P1' | 'P2' | 'P3' | 'P2' | Alert severity when the check fails |
Exposing a health endpoint
app.get('/health', async (req, res) => {
const status = await sdk.getHealthStatus();
res.status(status.status === 'healthy' ? 200 : 503).json(status);
});Response shape:
{
"status": "degraded",
"checks": [
{ "name": "database", "status": "healthy", "latencyMs": 4 },
{ "name": "redis", "status": "unhealthy", "latencyMs": 0, "error": "ECONNREFUSED" },
{ "name": "stripe", "status": "healthy", "latencyMs": 120 }
]
}Overall status logic:
| Status | Meaning |
|---|---|
healthy | All checks passed |
degraded | Some passed, some failed |
unhealthy | All checks failed |
Alerts
Alert rules define conditions that, when met, send notifications and optionally trigger the Trasys SRE agent.
Defining a rule
sdk.defineAlert({
id: 'high-error-rate',
name: 'High Error Rate',
description: 'Error rate exceeded 5% for 2 consecutive minutes',
condition: 'error_rate("payment-service") > 0.05',
windowSeconds: 120, // evaluate over a 2-minute window
evaluationIntervalSeconds: 30, // check every 30 seconds
severity: 'P1',
notify: {
channels: ['slack:#incidents', 'pagerduty'],
escalation: [
{ afterMinutes: 5, to: 'on-call-engineer', via: ['sms', 'phone'] },
],
},
environment: 'production',
cooldownSeconds: 900, // don't re-fire for 15 minutes after resolving
runbook: 'rb-payment-errors',
});Alert severity
| Severity | Notification | SRE agent |
|---|---|---|
P1 | SMS + phone call immediately | Claude Opus (extended thinking) |
P2 | Slack immediately; SMS after 5 min | Claude Sonnet |
P3 | Slack only | Claude Haiku |
See AI SRE Agent for exactly what the agent does with a firing alert, the full alert condition function reference, and how rules sync from your code to the backend.
Defining multiple rules
sdk.defineAlerts([
{
id: 'p99-latency',
name: 'P99 Latency Spike',
condition: 'p99_latency("api-gateway") > 2000',
severity: 'P2',
notify: { channels: ['slack:#performance'] },
},
{
id: 'ai-cost-overrun',
name: 'AI Daily Cost Limit',
condition: 'daily_cost() > 500',
severity: 'P3',
notify: { channels: ['slack:#engineering'] },
},
]);Alert condition reference
Service health
error_rate("service-name") > 0.05 // error rate > 5%
success_rate("service-name") < 0.95 // success rate < 95%
p50_latency("service-name") > 500 // median latency > 500ms
p95_latency("service-name") > 2000 // p95 latency > 2s
p99_latency("service-name") > 5000 // p99 latency > 5s
request_rate("service-name") < 10 // fewer than 10 req/s (drop detection)AI and agents
daily_cost() > 500 // total AI cost today > $500
daily_cost("agent-name") > 100 // cost for a specific agent
token_cost_per_run("agent-name") > 0.50 // cost per agent run > $0.50
tool_failure_rate("agent", "tool") > 0.10 // tool failure rate > 10%
agent_loop_detected("agent-name") // agent entered an infinite loopHealth checks
health_check_status("database") == "unhealthy"Combining conditions
p99_latency("checkout") > 3000 AND error_rate("checkout") > 0.02Notification channels
| Format | Example |
|---|---|
slack:#channel | 'slack:#incidents' |
slack:@user | 'slack:@on-call' |
email:address | 'email:alerts@company.com' |
sms | 'sms' |
phone | 'phone' |
pagerduty | 'pagerduty' |
Escalation
notify: {
channels: ['slack:#incidents'],
escalation: [
{ afterMinutes: 0, to: 'primary-oncall', via: ['sms'] },
{ afterMinutes: 10, to: 'secondary-oncall', via: ['sms', 'phone'] },
{ afterMinutes: 20, to: 'engineering-lead', via: ['phone'] },
],
},afterMinutes: 0 fires at the same time as the initial notification. Higher values escalate if the alert is not acknowledged within that time.
Next steps
- AI SRE Agent — what happens after an alert fires, and the full condition function reference
- TQL — query health check results over time with
FROM health_checks - Session Tracking — correlate health failures with specific user activity
Transport & Batching
How the SDK batches, sends, and retries telemetry — flush intervals, buffer size during network outages, and retry/backoff tuning.
AI SRE Agent
How Trasys's autonomous incident response works — alert severity maps to a specific Claude model, the SDK defines rules, and the backend investigates and triggers the agent.

