Alert Triage
Rank and deduplicate incoming alerts by severity, blast radius, and recent change correlation.
By OpenSRE
Prebuilt workflows your agent can run during investigations, from incident summaries to runbook retrieval.
Curated skill bundles for common SRE workflows.
Rank and deduplicate incoming alerts by severity, blast radius, and recent change correlation.
By OpenSRE
Cluster similar exceptions and stack traces to surface the highest-impact defects.
By OpenSRE
Pull service diagrams, ownership, and dependency maps for systems under investigation.
By OpenSRE
Recommend or trigger rollbacks when post-deploy error rates exceed safe thresholds.
By OpenSRE
Map affected services, dependencies, and customer impact from a single failure signal.
By OpenSRE
Parse CI logs to identify the failing step, flaky test, or infrastructure timeout.
By OpenSRE
Compare canary vs. baseline metrics to decide promote, hold, or rollback.
By OpenSRE
Summarize recent config, infra, and code changes near the incident start time.
By OpenSRE
Inspect pipeline failures, flaky tests, and blocked deploy gates across your CI system.
By OpenSRE
Correlate billing anomalies with recent deploys, autoscaling events, or config changes.
By OpenSRE
Determine which running services are affected by a newly disclosed vulnerability.
By OpenSRE
Read Grafana or Datadog dashboards and explain what the signals mean for the incident.
By OpenSRE
Summarize what changed in the latest release — config, images, migrations, and feature flags.
By OpenSRE
Follow request paths through microservices to locate latency and error hotspots.
By OpenSRE
Trace hostname resolution failures across internal and external DNS providers.
By OpenSRE
Find the right on-call rotation, manager chain, and vendor contacts for escalation.
By OpenSRE
Check flag states, targeting rules, and kill-switch readiness during incidents.
By OpenSRE
Review excessive permissions and recent policy changes tied to an access incident.
By OpenSRE
Condense active incidents into a timeline, impact summary, and next steps for responders.
By OpenSRE
Ask targeted questions to on-call engineers and synthesize answers into investigation context.
By OpenSRE
Search internal wikis and docs for prior incidents, fixes, and tribal knowledge.
By OpenSRE
Inspect target group health, TLS termination, and routing rule misconfigurations.
By OpenSRE
Cross-reference logs across services to find error patterns and request traces.
By OpenSRE
Spot unusual metric behavior against baselines and seasonal patterns.
By OpenSRE
Trace connectivity failures to misconfigured policies, DNS, or service mesh rules.
By OpenSRE
Evaluate node pressure, disk usage, and scheduling issues across the cluster.
By OpenSRE
Summarize open incidents, recent deploys, and known flaky systems for shift changes.
By OpenSRE
Inspect crash loops, OOM kills, probe failures, and resource limits for workloads.
By OpenSRE
Generate structured postmortem outlines with timeline, root cause, and action items.
By OpenSRE
Verify pre-deploy checks — tests, security scans, and approval policies — before promotion.
By OpenSRE
Walk through runbook steps, track progress, and record outcomes in the incident thread.
By OpenSRE
Find and surface the most relevant runbooks for the current alert or service.
By OpenSRE
Identify expired or soon-to-expire credentials referenced in failing services.
By OpenSRE
Diagnose error budget burn, identify violating SLIs, and suggest remediation.
By OpenSRE
Draft customer-facing incident updates from internal investigation notes.
By OpenSRE
Investigate failed uptime and browser checks with regional and endpoint breakdowns.
By OpenSRE
Compare live infrastructure against declared state and highlight unmanaged changes.
By OpenSRE
Summarize plan output, flag risky destroys, and explain dependency ordering.
By OpenSRE
Correlate blocked traffic patterns with false positives or active attack signatures.
By OpenSRE