Anomaly Root Cause
Anomaly Root Cause
Section titled “Anomaly Root Cause”Anomaly detection tells you that a service’s spend jumped. Root cause tells you why: which change happened before the rise, who made it, and on which resource. Every cause cites the evidence it is based on, so you can check it in one click rather than trust it.
What a root cause shows
Section titled “What a root cause shows”When an anomaly is detected, Xplorr reads the change history of the account the anomaly is in, ranks what it finds, and attaches the result to the anomaly before it is sent. A root cause has:
- A short explanation, two or three sentences, starting with the service,
region and daily increase, then the most likely cause, citing evidence as
[e1],[e2]. - A confidence level: High, Medium or No clear cause (see Confidence levels).
- Evidence: up to three change events, each with its time (UTC), who made it, the action, the resource, and a link to the event in your cloud console where the provider has one.
- Cost drivers (AWS): up to three usage types whose cost rose most on the spike day, each with the day’s cost against its 7 day average.
- Changes in the window: when nothing matches the spike, the write events that were found are listed instead, so you can see what did change.
Where the evidence comes from
Section titled “Where the evidence comes from”| Source | What is read | Used for |
|---|---|---|
| AWS CloudTrail | Management write events in the anomaly’s region, from CloudTrail event history. The service’s own events are read first (for example s3.amazonaws.com for an S3 spike), then every other write in the account | AWS accounts |
| Azure Activity Log | Administrative write and action operations on the subscription that succeeded | Azure subscriptions |
| GCP Cloud Audit Logs | Admin Activity entries for the project | GCP projects |
| Kubernetes changes | Worked out from the cluster cost data Xplorr already collects: a new namespace, a new controller, a controller that ran more pods, or one whose CPU, memory or GPU cost jumped | Clusters linked to the anomaly’s cloud account |
| AWS cost drivers | One Cost Explorer request for the anomaly’s service and region, by usage type, over the spike day and the 7 days before it | AWS accounts |
The audit logs are read on demand when an anomaly fires. Xplorr does not copy your audit logs. Only the cited evidence is stored with the anomaly: the actor, action, resource, time and a few values such as an instance type or a desired capacity. Full event payloads are never kept.
Frequent events that never change what you pay are dropped before anything is ranked: agent heartbeats (for example the Systems Manager agent), sign in and credential exchanges, parameter, secret and key reads, and log streams that services open on their own. Tagging and IAM changes are never cited as a cause.
Confidence levels
Section titled “Confidence levels”Each change event is scored out of 100 by fixed rules. The AI plays no part in the score.
| Part | Points | What earns them |
|---|---|---|
| Match | up to 40 | The same resource (40), the same service and region (25), or the same account (10) |
| Timing | up to 25 | Full points during the spike day or the 24 hours before it, falling to 5 at 48 hours before, and 0 after the day ends |
| Cost relevance | up to 20 | Adding capacity, such as a launch, scale out or larger size (20); changing how a resource bills, such as a storage class or memory setting (10) |
| Size fit | up to 10 | How well the size of the change matches the size of the rise |
| Your feedback | plus or minus 5 | Your organization’s past answers about this kind of event |
The confidence comes from the highest scoring event:
| Confidence | Top score | Meaning |
|---|---|---|
| High | 70 or more | A change that fits the spike on service, timing and type. Check the cited event first |
| Medium | 40 to 69 | A plausible change, with a weaker match on resource, timing or size |
| No clear cause | under 40, or no events | Nothing in the change history fits. For AWS, the cost drivers usually still show where the money went |
Even at High, the explanation says “most likely”. It is a lead to check, not a verdict.
Cost drivers for usage driven spikes
Section titled “Cost drivers for usage driven spikes”Many spikes have no change behind them: more requests, more data transferred,
more storage. For AWS, the cost drivers show which usage types rose, for
example DataTransfer-Out-Bytes or TimedStorage-ByteHrs, and by how much a
day. When one usage type accounts for at least half of the rise, the
explanation says “Most of the increase is …”, otherwise “The largest usage
increase is …”.
The breakdown is one Cost Explorer GetCostAndUsage request per anomaly, and
AWS bills each request (USD 0.01) to your own account. To keep that small, no
request is made when:
- the daily rise is under 5 (in the account’s billing currency), or
- the service that spiked is AWS Cost Explorer itself. Its cost is API requests, billed per request, made by the tools that read your bill (Xplorr’s syncs among them), and the explanation says so without asking for another breakdown.
Azure and GCP anomalies have no cost drivers yet. They are explained from the audit log and Kubernetes changes only.
How the explanation is written
Section titled “How the explanation is written”The rules above decide what the cause is and how confident Xplorr is. An AI model only turns the top evidence into two or three plain sentences, and what it writes is checked before it is kept:
- every citation such as
[e1]must point at real evidence, - it may not name a resource, person or number that is not in the evidence,
- it may not claim certainty.
An explanation that fails any check is replaced by a template sentence built from the top evidence item, for example “Amazon RDS in us-east-1 rose $38.40/day (+212%). Most likely cause: ops-deploy ran CreateDBInstance on reports-db at 2026-09-24 14:05 UTC [e1].”
The template is also used when the organization has no AI features (the Free plan) or has used its monthly AI budget. The evidence, confidence and cost drivers are the same either way.
Permissions per cloud
Section titled “Permissions per cloud”Root cause reads the audit log with the same credentials Xplorr already uses for the account. These permissions are part of the setup guides; if you connected an account before they were added, grant them now.
| Cloud | Permission | Notes |
|---|---|---|
| AWS | cloudtrail:LookupEvents | The AuditLog group of the AWS policy. Reads CloudTrail event history, which keeps 90 days of management events with no trail to set up |
| AWS | ce:GetCostAndUsage | For cost drivers. Already in the base AWS policy, since cost sync uses it |
| Azure | Reader or Monitoring Reader on the subscription | Carries Microsoft.Insights/eventtypes/values/read. Cost Management Reader alone cannot read the Activity Log. See Connect Azure |
| GCP | Logs Viewer (roles/logging.viewer) on the project | Carries logging.logEntries.list for Admin Activity audit logs, which are always on and free. See Connect GCP |
Kubernetes changes need nothing extra: they come from the cost data of a connected cluster.
Without the permissions
Section titled “Without the permissions”A missing permission never stops an anomaly from being sent. Instead:
- The explanation says what could not be read and what to grant, for example “Xplorr could not read the audit log: grant cloudtrail:LookupEvents to include it.”
- Kubernetes changes and, for AWS, cost drivers are still used.
- Every sync checks whether the audit log is readable. Alerts shows an amber banner listing each account whose audit log could not be read, with the permission to grant and a link to its setup guide. The account’s row in Cloud Accounts and its entry on Data health show the same warning. All of them clear after the next sync once the permission is in place.
In the console
Section titled “In the console”Open Alerts from the sidebar. The anomaly list has a Cause column with the confidence and the first sentence of the explanation. Click an anomaly to open its details:
- Cause: the confidence, the explanation, any missing permission, and the cost drivers. A Partial badge means a source was slow or throttled, so not every event was read.
- Evidence, or Changes in the window when nothing matched.
- Respond: Expected, Investigating or Wrong cause (see below).
- History: every answer given, by whom, and whether in the console or from Slack.
An anomaly with no root cause yet, such as one detected before root cause was available, has a Find cause button. It works for anomalies up to 90 days old, since that is as far back as CloudTrail event history goes, and usually takes under a minute.
Answering an anomaly
Section titled “Answering an anomaly”Admins and members can answer an anomaly. Viewers can read it.
| Answer | What it does |
|---|---|
| Expected | Marks the anomaly acknowledged, with an optional reason. You can also choose Don’t flag this again for 30 days for the same account, service and region |
| Investigating | Marks the anomaly investigating and assigns it to you |
| Wrong cause | Records that the cited cause is wrong, with an optional note on the real cause. The status does not change. Events of the same kind rank lower for your organization next time |
Wrong cause is only offered when a cause was cited.
In Slack and email
Section titled “In Slack and email”An anomaly with a root cause is sent once the cause has been found, usually a minute or two after detection.
Slack (your organization’s Slack webhook): the message shows the account, service and region, actual against baseline cost, the confidence, the explanation, the cost drivers and the evidence, with a link to each event. Expected, Investigating and Wrong cause open the anomaly in the console with that answer chosen, and View in Xplorr opens it without one.
Email (active admins who have not turned off anomaly alerts): a box headed Likely cause, high confidence (or medium), Where the increase came from when only cost drivers explain it, or No clear cause found, then the cost drivers and evidence, and a Was this right? line with the same three answers.
If finding the cause fails, the anomaly is still sent, in the usual layout, marked “Cause not determined.”
Limits
Section titled “Limits”- The window is 48 hours before the spike day to the end of that day (UTC). A change made earlier, whose cost built up slowly, is not found, and a change after the spike day is never cited.
- Management events only. CloudTrail data events (such as S3 object reads), Azure data plane operations and GCP Data Access logs are not read. Spikes driven by traffic or usage rather than by a change show up as cost drivers on AWS, or as no clear cause on Azure and GCP.
- Up to 500 events per read. In a very busy account the oldest writes in the window may not be read. The service’s own events are read first, so they are the last to be cut.
- Services Xplorr knows how to match. Events are matched for the main compute, database, storage, container, network, CDN and analytics services of each cloud (for example EC2, RDS, S3, EKS, Lambda, CloudFront, Virtual Machines, AKS, SQL Database, Compute Engine, GKE, Cloud SQL, BigQuery and Cloud Run). For another service the result is No clear cause, with the changes in the window listed.
- Kubernetes changes are inferred from daily cost data, not read from the Kubernetes API or its audit log, so they show that a controller scaled or appeared, not who changed it.
- Cost drivers are AWS only.
Does root cause change how anomalies are detected? No. Detection is the same; root cause adds an explanation to each anomaly it finds.
Does it write anything to my cloud accounts? No. Every call is a read with the credentials you already gave Xplorr.
What does it cost me? Reading CloudTrail event history, the Azure Activity Log and GCP Admin Activity logs is free. Each AWS cost driver breakdown is one Cost Explorer request, USD 0.01, and only for rises of 5 a day or more. The AI explanation counts toward your organization’s monthly AI budget.
Why does an anomaly say “No clear cause”? No change in the window scored 40 or more. Check the cost drivers and the changes in the window, and whether the audit log permission is missing for the account.