RegencyOps
🏠 Home ℹ️ About Us 📦 Products 🖥️ Enterprise 🛰️ Stratos – Alert2RootCause 🎯 Role Simulation 🎓 Education
🛰️ Stratos · Incident AI Command Center

From Grafana Alert to
Root Cause — in Minutes.

Stratos reads your monitoring alert, walks it through your mapped business flow, and identifies the closest impacted component with 99% accuracy — highlighting the affected path in red, from the data layer to the application layer to infrastructure.

Start 3-Day Free Trial → ▶ See How It Works
Stratos incident command center dashboard
The Problem

Alerts tell you something broke. Not what — or why.

Every incident starts the same way: a wall of alerts, a scramble across dashboards, and an engineer manually tracing which service, which app, and which customer are actually affected.

⚠ Without Stratos

  • An alert fires — but nobody knows which business application it actually impacts until someone manually traces upstream/downstream dependencies.
  • Root-cause analysis is manual: paging through logs, guessing at a 5-Whys, searching old tickets for "have we seen this before?"
  • MTTR balloons while the on-call engineer correlates data across Grafana, Jira, ServiceNow, and their own memory.
  • By the time impact is understood, SLA clocks are already blown.

✓ With Stratos

  • Stratos ingests the alert and automatically walks your pre-mapped business flow — infrastructure → services/components → business applications.
  • It finds the closest impacted component and confirms the match against monitoring, logs, feature-flag changes, and prior-incident knowledge — with 99% detection accuracy.
  • The impacted path is highlighted in red, live, across every layer — from data to application to infra — so the right team is looped in immediately.
  • A 5-Whys, Fishbone analysis, and postmortem are generated automatically once the incident resolves.
How It Works

One alert. A fully mapped blast radius.

Stratos connects to Grafana (and your ticketing, logging, and feature-flag tools) once. From then on, every alert is automatically correlated against the business flow you've mapped — no manual triage.

1 · Alert
PaymentSvcErrorRate 97.4%
2 · Component
db-payment pool
3 · Business App
Checkout Service
4 · Downstream Risk
Billing, CRM
5 · Root Cause
99% confidence
▶ Demo video — drop into /enterprise/AICommandCare/vdo/stratos-how-it-works.mp4
Features

Everything from alert intake to postmortem.

📊 Grafana-Integrated Alert Intake

Connect your Grafana service account once — Stratos pulls live and historical alerts automatically, no manual forwarding.

🗺️ Business Flow Mapping

Model your instances, architecture, and business applications — with tier, upstream/downstream, and owning teams — once, and reuse for every incident.

🎯 99% Accuracy Root-Cause ID

Deterministic correlation plus AI-assisted analysis cross-checks logs, deploys, feature flags, and past-incident knowledge before naming a root cause.

🔴 Red-Highlighted Impact Path

The full blast radius — from the failing component up through business apps to infra — lights up in red so nobody has to ask "is this affected too?"

⏱️ SLA & Severity Matrix

Tiered SLA timers (ack/resolution) run automatically per app criticality, so escalation never depends on someone remembering the runbook.

📝 Auto-Generated Postmortems

5-Whys, Fishbone analysis, sequence of events, and an action plan — compiled the moment an incident is marked resolved.

Why It Matters

Built for the outcomes that matter to your team.

Not just another dashboard — Stratos is measured by what it removes from your incident process.

01

RCA in Plain English, Under 5 Minutes

Turns an alert into a root-cause explanation any stakeholder can understand — not just the engineer who's paged.

02

Correlates Every Monitoring Alert

Pulls signals from all your monitoring tools into one unified view instead of siloed, per-tool dashboards.

03

Reduces MTTR

Cuts mean-time-to-resolution by skipping the manual triage scramble across logs, tickets and tribal knowledge.

04

Meets Your SLAs

Faster detection and root-cause identification means fewer breached SLA windows.

05

Fewer Layer Hops

Cuts the L1 → L2 → L3 escalation chain by getting the right team engaged from the very first alert.

Documentation

Guide — set it up as an Admin, use it as a Viewer.

Pick the guide that matches what you're here to do. Setup is a one-time job for one admin — everyone else just reads incidents.

This is the setup reference for whoever connects tools and configures Stratos — a one-time job, usually done by one admin. Everyone else on the team only needs the Viewer guide.

Part A — Connect Your Tools

Stratos pulls data from your monitoring, ticketing, code and knowledge tools, maps it onto your business flow, and correlates it down to a single most-likely root cause. Connect tools in this order:

#CategoryWhat it's for
1Monitoring (Grafana)Live alerts — the primary signal an incident is happening.
2Ticketing (Jira / ServiceNow)Tickets as an alternate trigger; auto-creates/updates tickets during an incident.
3CI/CD (GitHub)Recent commits/deploys as extra root-cause evidence — "deployed 12 min before the incident."
4Knowledge (Confluence / Docs / KB)Past-incident history and matched runbook steps.
5Layer Configuration (Part B, below)Your business stages, dependencies, apps, on-call, SLAs.
6On-call / Comms / Logs / InfraPaging, Slack posting, log search, CMDB.

Tool-by-tool setup

📡 Monitoring — Grafana
Why: the primary trigger — live firing alerts are what a real incident looks like.
  1. In Grafana: Administration → Users and access → Service accountsAdd service account, role Viewer (read-only).
  2. Open it → Add service account token → copy the token (shown once).
  3. In Stratos Configuration → Grafana card → paste the URL and token.
  4. On Grafana Cloud, also fill the CORS proxy field — Grafana Cloud doesn't allow direct browser calls otherwise.
  5. Test connection, then Extract dashboards & alerts.
Read-only — Stratos never writes to Grafana.
🎫 Ticketing — Jira
Why: lets a support ticket trigger an incident, and posts runbook comments back to it.
  1. id.atlassian.com → Security → API tokens → Create API token (classic, not "with scopes").
  2. In Stratos: fill Jira URL, account email, the token, and your Project Key (e.g. OPS).
  3. Fill the CORS Proxy field — Atlassian blocks direct browser calls too.
  4. Test connection, then Save Configuration.
Read-only by default; auto-ticket-creation is a separate opt-in checkbox.
🔀 Code Repository & CI/CD — GitHub
Why: a commit or deploy shortly before an incident is one of the strongest root-cause signals.
  1. GitHub → Settings → Developer settings → Fine-grained personal access tokens → generate one scoped to only the repo that backs the affected service.
  2. Permissions → Contents: Read-only (optionally Actions: Read-only too).
  3. In Stratos: fill Org/Owner, Repository, and the token. No CORS proxy needed here.
💬 Communications — Slack
Why: post incident updates straight to your team's channel.

One-click "Connect with Slack" (OAuth) is available, or skip it entirely and paste a Bot Token directly — both work.

📚 Knowledge — Confluence / Docs / Jira history
Why: feeds "similar past incident" matching and matched remediation runbook steps.

Past Jira tickets and Confluence runbooks import from the Docs / Links / API card. A pasted webpage or Google Doc link works too — a raw PDF link doesn't (paste its text instead).

✨ AI Providers — Gemini / ChatGPT / Claude
These are configured centrally, not per-install — one admin sets them once in the admin panel and every Stratos user on the account picks them up automatically, no per-browser setup or redeploy needed.

Part B — Configure Your Business Flow (Layer Configuration)

This is what makes a root cause map back to an actual business step, not just a server name. Layers are defined once and reused for every incident:

LayerWhat you define
L0 · Executive SummaryAuto-generated — no setup.
L1 · Monitoring SignalsAuto — pulled from connected monitoring tools.
L2 · Business Process StagesYour customer journey (e.g. login → cart → payment → fulfilment), mapped to technical services.
L3 · Service Dependency GraphWho depends on whom — drives "origin node vs. symptom" reasoning.
L4 · Business ApplicationsYour app inventory.
L5 · Technical ArchitectureServices, components, how they connect.
L6 · InfrastructureAuto — pulled from connected infra tools.
L7 · User + Financial ImpactSLA minutes, revenue/min, affected user counts per stage.
L8 · Actions / RunbooksThe fix steps Stratos recommends per known failure pattern.
L9 · External DependenciesThird-party services you depend on.

Once Parts A and B are done, every future alert automatically walks this flow — no per-incident setup.

Versions & Pricing

Pilot free. Scale when you're ready.

Every plan is read-only by design — Stratos reads your monitoring, ticketing, and knowledge systems to find root cause. It never runs commands on your systems.

Pilot

Try it with your team, read-only
Free
for 3 days
  • Up to 5 seats
  • 1 business flow mapped
  • Grafana integration
  • Read-only — no config changes
Start Free Trial

Enterprise

Org-wide, custom SLAs & support
Contact Us
 
  • Unlimited seats
  • Custom SLA & severity matrix
  • Dedicated onboarding
  • Priority support
Free Trial

Start your 3-day free Pilot trial.

No credit card. We'll create your access code instantly — sign in at the Stratos login page with the code and your work email.

The Pilot trial is a single-seat, read-only preview for you as admin. Upgrade to Pro or Enterprise to add your team.