Most platforms ship AI agents the same week they write them. The agent prompt compiles, the demo runs, the feature goes out. Nobody measures whether the agent is safe on real customer data until a customer complains.
XIntelliSync makes that impossible. Every one of the 150+ launch agents passes through XAVS — the XIntelliSync Agent Verification System — before it reaches a production tenant. Not as a label. As a gate. The gate is ten scoring dimensions, nine of them hard pass-or-fail, and a deploy-tier tag that has to be earned on every major platform change.
V1 is not a sticker on an agent. V1 is the tier every agent must earn before it sees your real business — and the tier it loses the moment Horizon spots drift.
Six scoring lanes
Six lanes. Ten dimensions. Nine hard gates. One threshold.
Accuracy, calibration, safety, output integrity, regulatory currency, data contract consistency — and four more dimensions that sit underneath them. Every agent scores on all ten. Nine of them are pass-or-fail individually. The agent has to clear every hard gate on its own, not just on a weighted average.
Accuracy.
Does the agent actually solve the task it claims to solve — on structured scenarios, adversarial scenarios, and replays of real user data? Sampled across scenario mix, not one lucky run.
Gate rule: Hard gate. Agent cannot promote to V1 until accuracy clears the threshold on the scenario pack.
Safety.
Customer-harm potential — can the agent take an action that damages your books, your reputation, or your compliance posture? Includes escalation correctness: does the agent hand off to you when the right answer is "I don’t know" instead of guessing?
Gate rule: Hard gate. Customer-harm, escalation, and approval-required obedience are individually scored, individually pass-or-fail.
Compliance.
Does the agent follow the policy it was deployed under — tier gates, tenant scope, disclaimer presence, and the specific regulatory framework its action falls under? STP Phase 2. BAS. Fair Work awards. Privacy Act APPs.
Gate rule: Hard gate. Policy compliance, disclaimer presence, and regulatory currency awareness all scored individually.
Output integrity.
Can the agent show its work? Every number cited to a source. Every recommendation traceable to a rule, a record, or a framework. No black-box answers with no evidence trail.
Gate rule: Scored on evidence quality + output consistency. Low scores cap the agent at assist mode, not autonomous.
Operational reliability.
Tool-use correctness. Does the agent call the right tool with the right arguments in the right order? Latency reliability. Does it finish inside its budget? A slow agent is not a safe agent on a live OS.
Gate rule: Combined score on tool use + latency. Threshold calibrated per agent class, not one global bar.
Data hygiene.
PII redaction in all outbound surfaces. Prompt-injection resistance on user-supplied inputs. Your customer’s name does not leak to a log. An email saying "ignore previous instructions" does not flip the agent into a different personality.
Gate rule: Hard gate. PII redaction hygiene and prompt injection resistance are both pass-or-fail individually.
Six public lanes, ten underlying dimensions. The remaining four dimensions stay internal — not because they’re secret, but because scoring mechanics should be reviewed by engineers, not customers. The outcomes are public. The scoring surface is engineering.
Three trust tiers
V0 → V1 → V2+. Each tier earned. Each tier reversible.
Agents do not default to autonomous. They start in sandbox. They earn V1. They earn progressive autonomy one step at a time. A single drift event drops the agent back one tier — not one notch, one tier — until re-certification.
V0 · Sandbox.
The agent exists. It has not earned live data.
Access: Runs on synthetic scenarios and anonymised replays only. Never touches your real books, customers, payroll, or financial records.
Earns promotion by: Clears the full 10+ dimension scoring pack at the V1 threshold. Passes all nine hard gates individually. Only then does promotion happen.
V1 · Live-data eligible.
The agent has earned access to your real data.
Access: Runs on your live business data under Horizon supervision. Delivers outputs to your dashboards, your reports, your decisions. Every action logged.
Earns promotion by: Sustained performance under Horizon. Low drift. Zero rejected actions on a sample window. Accumulated trust across a configured number of executions.
V2+ · Progressive autonomy.
The agent has earned write actions and eventually autopilot.
Access: Moves through supervised → controlled autopilot → certified autonomous. Each step is earned, never defaulted, and always reversible.
Earns promotion by: Continuous XAVS re-verification plus Horizon runtime signal. A single drift event drops the agent back one tier until re-certification.
The default is doubt. V1 is the gate that earns trust.
Generic AI platforms treat trust as marketing. They ship the model, publish benchmark numbers, and hand drift detection to you — the customer who finds out at BAS time that the AI was wrong three weeks ago. XIntelliSync treats trust as code. Trust is the gate at the door (XAVS), the supervisor on the floor (Horizon), the ledger at the end of the day (XGVS).
V1 is what the gate checks. Ten dimensions. Nine hard gates. A threshold that gets stricter on higher-risk agent classes, never looser. Tax and payroll agents are permanently capped at supervised mode — Fair Work and ATO compliance require your approval on every action, always, no exceptions, no matter how many V2+ scores the agent accumulates.
Verified AI beats unverified AI on a financial platform. Not by a little. By the difference between "the numbers stay right" and "we’ll find out in a quarter when the audit opens."
Key takeaways
- V1 means ten dimensions scored at the certification threshold AND nine hard gates passed individually — not one weighted average.
- Six public scoring lanes: accuracy, safety, compliance, output integrity, operational reliability, and data hygiene.
- Three trust tiers: V0 (sandbox), V1 (live-data eligible), V2+ (progressive autonomy earned through Horizon runtime signal).
- Every major platform change triggers re-verification. Drift detected in production flags the agent for early re-verification.
- Every agent is subject to XAVS. Certified autonomous is earned, never defaulted. Tax and payroll agents are permanently capped at supervised.
V1 certification — questions answered.
What does V1 certified actually mean in practice?+
How often do agents re-verify?+
What happens when an agent fails V1?+
What’s the difference between XAVS and Horizon?+
Can I see the V1 score for a specific agent?+
Is V1 the same threshold for every agent?+
Do any agents skip V1 verification?+
How does V1 compare to generic AI platforms?+
Keep reading
More of the Trust Stack.
Pillar page
AI Verification — the full overview
All four Trust Stack systems — Omega, Horizon, XAVS, XGVS — and how they compose into a single verification engine.
Read pillarSpoke · Runtime sibling
Why AI Drift Compounds Silently on Live Data
How Horizon catches per-phase drift, calibration slippage, and distribution shift the moment they happen — before your books show it.
Read spokeSpoke · Risk sibling
Why AI Hallucinations Break Financial Data
Three real hallucination patterns that cost Australian SMBs at BAS, payroll, and audit — and how the Trust Stack catches them.
Read spokeSpoke · Compliance sibling
The 34+ Compliance Frameworks Every AI Action Must Pass
356+ gates across 6 stages covering ATO DSP, STP Phase 2, Fair Work, Privacy Act, Essential 8, SOC 2 and more.
Read spoke150+ agents. Ten dimensions. One gate each has to pass.
From $97/month AUD. XAVS verification runs on every agent on every tier. Not an add-on. Not an upgrade. The gate at the door. Built in Australia. Built for what’s next.