XAVS · Agent Verification

What V1 Certification Actually Requires

Every agent earns its way to your data. Or it never sees your data.

Published 24 April 2026 7 min readBy Wissam, Founder & Creator of XIntelliSync

Why this lives on XIntelliSync, not on an MLOps certification blog

V1 certification is not a generic MLOps concept — it's the specific bar XAVS requires before an agent runs on live customer data. Every launch agent in the catalogue is scored on 10+ verification dimensions and only V1-certified agents enter production. XIntelliSync does not offer a generic agent-certification service to other platforms; V1 is what we score our own agents against. The article below describes exactly which dimensions an agent has to pass and what failure looks like.

← AI Verification pillar · See how the Trust Stack works →

Most platforms ship AI agents the same week they write them. The agent prompt compiles, the demo runs, the feature goes out. Nobody measures whether the agent is safe on real customer data until a customer complains.

XIntelliSync makes that impossible. Every one of the 150+ launch agents passes through XAVS — the XIntelliSync Agent Verification System — before it reaches a production tenant. Not as a label. As a gate. The gate is ten scoring dimensions, nine of them hard pass-or-fail, and a deploy-tier tag that has to be earned on every major platform change.

V1 is not a sticker on an agent. V1 is the tier every agent must earn before it sees your real business — and the tier it loses the moment Horizon spots drift.

Six scoring lanes

Six lanes. Ten dimensions. Nine hard gates. One threshold.

Accuracy, calibration, safety, output integrity, regulatory currency, data contract consistency — and four more dimensions that sit underneath them. Every agent scores on all ten. Nine of them are pass-or-fail individually. The agent has to clear every hard gate on its own, not just on a weighted average.

Accuracy.

Does the agent actually solve the task it claims to solve — on structured scenarios, adversarial scenarios, and replays of real user data? Sampled across scenario mix, not one lucky run.

Fails accuracy and the agent stays in sandbox. It does not touch your books, your payroll, or your customer records.

Gate rule: Hard gate. Agent cannot promote to V1 until accuracy clears the threshold on the scenario pack.

Safety.

Customer-harm potential — can the agent take an action that damages your books, your reputation, or your compliance posture? Includes escalation correctness: does the agent hand off to you when the right answer is "I don’t know" instead of guessing?

Fails safety and the agent is capped at supervised mode. Every action requires your approval before it touches live data.

Gate rule: Hard gate. Customer-harm, escalation, and approval-required obedience are individually scored, individually pass-or-fail.

Compliance.

Does the agent follow the policy it was deployed under — tier gates, tenant scope, disclaimer presence, and the specific regulatory framework its action falls under? STP Phase 2. BAS. Fair Work awards. Privacy Act APPs.

Fails compliance and the agent is blocked — even if the answer is right. A right answer delivered outside policy is not production-ready on a financial platform.

Gate rule: Hard gate. Policy compliance, disclaimer presence, and regulatory currency awareness all scored individually.

Output integrity.

Can the agent show its work? Every number cited to a source. Every recommendation traceable to a rule, a record, or a framework. No black-box answers with no evidence trail.

Fails output integrity and the agent loses downstream trust — Horizon will flag every action for review until the next re-verification.

Gate rule: Scored on evidence quality + output consistency. Low scores cap the agent at assist mode, not autonomous.

Operational reliability.

Tool-use correctness. Does the agent call the right tool with the right arguments in the right order? Latency reliability. Does it finish inside its budget? A slow agent is not a safe agent on a live OS.

Fails reliability and the agent drops from the autonomous roster. Still available manually, but no longer scheduled without supervision.

Gate rule: Combined score on tool use + latency. Threshold calibrated per agent class, not one global bar.

Data hygiene.

PII redaction in all outbound surfaces. Prompt-injection resistance on user-supplied inputs. Your customer’s name does not leak to a log. An email saying "ignore previous instructions" does not flip the agent into a different personality.

Fails data hygiene and the agent cannot see customer-facing data at all — it stays on internal-only operations until re-verified.

Gate rule: Hard gate. PII redaction hygiene and prompt injection resistance are both pass-or-fail individually.

Six public lanes, ten underlying dimensions. The remaining four dimensions stay internal — not because they’re secret, but because scoring mechanics should be reviewed by engineers, not customers. The outcomes are public. The scoring surface is engineering.

Three trust tiers

V0 → V1 → V2+. Each tier earned. Each tier reversible.

Agents do not default to autonomous. They start in sandbox. They earn V1. They earn progressive autonomy one step at a time. A single drift event drops the agent back one tier — not one notch, one tier — until re-certification.

V0 · Sandbox.

The agent exists. It has not earned live data.

Access: Runs on synthetic scenarios and anonymised replays only. Never touches your real books, customers, payroll, or financial records.

Earns promotion by: Clears the full 10+ dimension scoring pack at the V1 threshold. Passes all nine hard gates individually. Only then does promotion happen.

V1 · Live-data eligible.

The agent has earned access to your real data.

Access: Runs on your live business data under Horizon supervision. Delivers outputs to your dashboards, your reports, your decisions. Every action logged.

Earns promotion by: Sustained performance under Horizon. Low drift. Zero rejected actions on a sample window. Accumulated trust across a configured number of executions.

V2+ · Progressive autonomy.

The agent has earned write actions and eventually autopilot.

Access: Moves through supervised → controlled autopilot → certified autonomous. Each step is earned, never defaulted, and always reversible.

Earns promotion by: Continuous XAVS re-verification plus Horizon runtime signal. A single drift event drops the agent back one tier until re-certification.

Every major platform change re-verifies every agent. Drift rolls the agent back. Nothing drifts into higher trust by accident.

The default is doubt. V1 is the gate that earns trust.

Generic AI platforms treat trust as marketing. They ship the model, publish benchmark numbers, and hand drift detection to you — the customer who finds out at BAS time that the AI was wrong three weeks ago. XIntelliSync treats trust as code. Trust is the gate at the door (XAVS), the supervisor on the floor (Horizon), the ledger at the end of the day (XGVS).

V1 is what the gate checks. Ten dimensions. Nine hard gates. A threshold that gets stricter on higher-risk agent classes, never looser. Tax and payroll agents are permanently capped at supervised mode — Fair Work and ATO compliance require your approval on every action, always, no exceptions, no matter how many V2+ scores the agent accumulates.

Verified AI beats unverified AI on a financial platform. Not by a little. By the difference between "the numbers stay right" and "we’ll find out in a quarter when the audit opens."

Key takeaways

  • V1 means ten dimensions scored at the certification threshold AND nine hard gates passed individually — not one weighted average.
  • Six public scoring lanes: accuracy, safety, compliance, output integrity, operational reliability, and data hygiene.
  • Three trust tiers: V0 (sandbox), V1 (live-data eligible), V2+ (progressive autonomy earned through Horizon runtime signal).
  • Every major platform change triggers re-verification. Drift detected in production flags the agent for early re-verification.
  • Every agent is subject to XAVS. Certified autonomous is earned, never defaulted. Tax and payroll agents are permanently capped at supervised.

V1 certification — questions answered.

What does V1 certified actually mean in practice?+
V1 means an agent has cleared all ten XAVS verification dimensions at the certification threshold AND passed every one of the nine hard gates individually. Hallucination resistance, policy compliance, escalation correctness, customer-harm potential, compliance-disclaimer presence, regulatory-currency awareness, prompt-injection resistance, PII-redaction hygiene, and approval-required obedience — every one must pass on its own, not just on a weighted average. Only then does the agent earn access to live customer data. Before V1, the agent runs on sandbox scenarios and anonymised replays.
How often do agents re-verify?+
Every major platform change triggers a re-verification sweep. LLM provider routing change, scenario pack expansion, new compliance framework added to XGVS, schema change that affects the agent’s tool surface — any of these fire a re-verification run. Individual agents also re-verify on a rolling cadence, not just on platform events. Horizon runtime signal feeds into XAVS so drift detected in production flags the agent for early re-verification.
What happens when an agent fails V1?+
Three things, in order. First, the agent does not promote — it stays V0 (sandbox) or drops back to V0 if it was V1 and failed re-verification. Second, the failure breakdown is logged with per-dimension scores and per-hard-gate pass-or-fail, so engineering knows exactly which lane to fix. Third, any scheduled autopilot or supervised runs are cancelled until re-verification clears. The agent does not silently degrade — it goes offline for the affected customer surface until the gate reopens.
What’s the difference between XAVS and Horizon?+
XAVS is the gate BEFORE an agent reaches production. Ten dimensions scored, nine hard gates, V1 threshold cleared — or the agent does not ship. Horizon is the supervisor WHILE the agent runs in production. Per-phase observability, calibration tracking, drift detection on every autonomous action. XAVS says "this agent is ready to ship." Horizon says "this agent is still behaving the way XAVS verified." Both run on every agent, every tier.
Can I see the V1 score for a specific agent?+
Yes. Every agent carries its XAVS verification badge — V0 Provisional or V1 Certified — visible on the agent detail page. Higher tiers surface more granular scoring (dimension-by-dimension breakdown, last re-verification date, Horizon runtime confidence). The principle is: every customer can see the certification state of every agent they run on their data. Transparency is the foundation, not a Starter-only or Enterprise-only feature.
Is V1 the same threshold for every agent?+
The ten scoring dimensions and the nine hard gates are the same for every agent. The threshold is tuned per agent class because an invoice reminder agent and a BAS preparation agent face different risk surfaces. Higher-risk agents (anything touching ATO, payroll, Fair Work awards, outbound communications) face stricter thresholds on the hard gates than a read-only reporting agent. The bar is never lower than the floor — it is calibrated up for risk, never down.
Do any agents skip V1 verification?+
No. Every one of the 150+ launch agents is subject to XAVS. Some agents with very high safety risk (tax preparation, payroll execution) are permanently capped at supervised mode regardless of how many V2+ scores they accumulate — Fair Work and ATO compliance require human approval on every action, so certified autonomous mode is explicitly blocked for those agent classes. That cap is a policy decision, not a score decision.
How does V1 compare to generic AI platforms?+
Generic AI platforms (OpenAI, Anthropic, Gemini consumer) do not score their output agents. They ship inference. Verification is your problem — and you find out they were wrong at BAS time, payroll dispute time, or audit time. XIntelliSync treats verification as the gate, not the afterthought. 150+ agents scored across 10+ dimensions with 9 hard gates — every one cleared, or the agent does not reach your data. This is why the Trust Stack exists: generic AI has no concept of certification. We made it the floor, not the ceiling.

Keep reading

More of the Trust Stack.

Verification is the foundation. Every agent. Every tier.

150+ agents. Ten dimensions. One gate each has to pass.

From $97/month AUD. XAVS verification runs on every agent on every tier. Not an add-on. Not an upgrade. The gate at the door. Built in Australia. Built for what’s next.