The CAI Score

The CAI Score: governance you can put a number on.

How ColleagueAI classifies agent autonomy, maps controls to risk, and produces the evidence your auditors and regulators expect.

Enterprise architecture

Designed to run in your tenant, not in a shared AI data lake.

Agent runtime and customer business content remain in the customer-controlled Microsoft environment. ColleagueAI supplies the agent blueprint, the CAI Score governance pattern, deployment documentation and versioned package updates.

Customer data plane

The runtime, identity checks, retrieval and audit trail all run inside your environment.

Entra ID / Active Directory
User identity, groups, conditional access and permissions.
M365, SharePoint, Teams, Outlook
Business content remains under customer tenant controls.
Azure OpenAI or approved model endpoint
Inference path selected by the customer deployment model.

ColleagueAI package layer

We package the agent blueprint and its controls so your teams do not start from a blank prompt, a blank workflow, or a blank audit model.

Deployable blueprint
Workflow, prompts, controls and deployment pattern.
CAI Score classification
Autonomy tier, human oversight model and risk notes.
No customer data by default
Package delivery does not require ColleagueAI to process customer business content.

Permission-aware retrieval

RAG patterns should inherit the user’s existing Entra ID and M365 permissions. The agent should only retrieve or summarize content the user is already authorized to access.

Human oversight by tier

L3 and L4 packages are designed with review, approval, escalation and audit evidence patterns.

Use-case dependent risk

AI Act and regulated-use classification depends on customer process, data and deployment context. This is governance evidence, not a legal compliance guarantee.

Internal build from scratch

  • Discovery, prompt design and controls start from zero.
  • Security review repeats for each use case.
  • Audit logs and oversight are often added late.
  • Maintenance remains fully internal.

ColleagueAI package

  • Deployable blueprint and governance model are structured up front.
  • CAI Score defines autonomy, oversight and risk posture.
  • Designed for pilot deployment in weeks, subject to environment readiness.
  • Evidence and handover materials are part of the package design.
The philosophy

Most AI projects fail on governance, not capability. So we put the blueprint first.

ColleagueAI is the trust layer for enterprise AI — the governance framework your classified agents run inside. The hard part of enterprise AI isn’t building an agent that works in a demo, it’s governing the sprawl once dozens of them are live: knowing where each one is allowed to act, who’s accountable when it does, and being able to prove it to an auditor. Colleague AI gives you that framework first (the governance blueprint) then fits classified agents into it at the right tier.

01

Blueprint before bots

The blueprint comes first: every step of your process gets the CAI tier it's allowed to operate at, and each slot is filled only by an agent classified at or below that tier. Agents fill slots in a governed design, they don't get bolted onto a process and hope for the best.

02

Safe by tier, not by promise

An agent's CAI tier is a hard contract about its autonomy. Low-tier agents only draft and suggest. High-tier agents assist a named human who stays accountable. Safety is structural, not a marketing claim.

03

It runs in your house

Agents run inside your own Microsoft and cloud environment. Colleague AI hosts the governance layer, metadata, configuration, scores, audit trail, entitlements, and partner-registration records. In production deployments, customer operational data is designed to remain inside the customer tenancy. Public website demos, forms, partner registrations, and other submitted website inputs are separate ColleagueAI-hosted flows and should not be used for confidential production data.

Why trust us

Trust should be demonstrated, not requested.

We prefer evidence over claims: defined tiers, named accountability, operating controls, and audit records that can be reviewed.

01

Built by operators, not just model-builders

Colleague AI is built by people who have run regulated finance, insurance and SAP operations, and sat through the audits that follow. The CAI Score encodes how those functions actually reason about risk and control, not how a demo looks.

02

A documented method, not a black box

The CAI Score is an open framework: defined tiers, named human roles, logging and review periods. You can see exactly how every agent is classified, and push back if you disagree. A methodology that holds up to scrutiny is one you can actually stand behind.

03

Verify the evidence

Agents run in your tenant. Customer business content remains in the customer-controlled Microsoft environment, while ColleagueAI processes separate governance-layer metadata as described in the Trust Center. Actions are logged and attributable, and the evidence is designed to support governance and compliance review.

The CAI Score

Five levels of autonomy. Clear accountability at each level.

The higher the tier, the more an agent can do. The controls also increase: human oversight, logging, review points, and named accountability.

L1

Assist

Informs and answers. Takes no action and changes no record. The human does the work; the agent just makes it faster to find and understand.

Human role
Does everything
L2

Draft

Produces a work product (a document, a query, a report, an outreach message) for a human to review and approve. Nothing the agent makes is used until a person signs off.

Human role
Reviews & approves
L3

Operate

Executes routine, low-risk actions inside a bounded workflow, classify, route, fulfil, log. Exceptions and anything unusual are handed to a human. Every action is time-stamped.

Human role
Owns exceptions
L4

Decide (supervised)

Supports decisions and controls in higher-stakes processes, compliance, contracts, security, four-eyes. A named human remains accountable for the call; the agent assists and evidences it. Built for high-risk-process scrutiny.

Human role
Stays accountable
L5

Autonomous

Acts independently within hard, pre-approved guardrails. Reserved for the highest level of classification, and not used by any agent in this catalogue today.

Human role
Sets the rails

// This catalogue ships at L2-L4. L4 agents require senior sign-off before go-live.

CAI Score × EU AI Act

How CAI tiers support EU AI Act assessment.

The AI Act classifies high-risk systems by intended purpose and the Article 6 criteria, not by CAI tier. For Annex III use cases, the CAI Score can structure the governance and evidence assessment, but it does not determine legal classification.

CAI tiers structure governance controls; the AI Act determines legal roles and obligations. The Act splits duties between the provider of a system (whoever develops it and puts it on the market under their own name) and the deployer (whoever runs it under their own authority). In a typical deployment Colleague AI is the provider of the agent package and you are the deployer — but under Article 25 you become a provider yourself if you rebrand it, modify it substantially, or repurpose it into a new high-risk use. The deployer's duties stay with you either way: real human oversight, keeping the agent's automatic logs at least six months, monitoring and reporting incidents, and informing affected people where required. The CAI Score gives both sides the evidence for these duties; it doesn't transfer or discharge them.

L1

Minimal risk

Outside Annex III, L1 carries no AI Act obligations beyond existing law. Article 50's disclosure duty (tell the person they're talking to AI) applies only if the tool is customer- or public-facing, not to an internal lookup or drafting aid.

Key article(s)
Art. 50, if public-facing
L2

Possible Art. 6(3) exception

An Annex III system may fall outside high-risk status under Article 6(3) where it does not pose a significant risk, including by not materially influencing decision-making, and at least one statutory condition applies, such as a narrow procedural or preparatory task. Profiling of natural persons remains high-risk. Any reliance on Article 6(3) must be documented before placing the system on the market or putting it into service.

Key article(s)
Art. 6(3) + 6(4)
L3

Higher-control deployment

An L3 agent used in an Annex III context may be high-risk depending on its intended purpose and the Article 6 criteria. Where the system is high-risk, applicable requirements can include risk management, logging and effective human oversight.

Key article(s)
Art. 9, 12, 14
L4

High-stakes, supervised

L4 is designed for high-stakes use cases where stronger oversight and documentation are appropriate. Where a system is high-risk under the AI Act, applicable requirements may include Article 14 human oversight, technical documentation and, where required, conformity assessment.

Key article(s)
Art. 14, 43, Annex IV
L5

Reserved, unused

Not because Article 5 bans it — unbounded autonomy isn't on that list. The problem is Article 14: a high-risk system no one can meaningfully oversee, override, or stop isn't compliant. That's the actual reason this tier stays reserved and unused in the catalogue today.

Key article(s)
Art. 14, unsatisfied

// Indicative mapping, not legal advice. Classification depends on the specific workflow, not the tier alone — always confirm Annex III scope and any Article 6(3) reliance, including the profiling exclusion, with counsel before go-live.

Deployment & safety

Your environment. Your data. Clear governance around the agent.

Agents run where your business data already lives. ColleagueAI holds the governance metadata, scoring, audit evidence, and token-economy indicators around them.

// architecture
Client-hosted by design

Agents run inside your Microsoft Copilot Studio, Power Automate and Azure estate. ColleagueAI hosts only the governance layer, scores, policies and audit metadata. No customer business data is processed by the ColleagueAI governance layer. Your business data remains in your Microsoft environment.

// accountability
A human is always named

The CAI tier defines exactly where a person stays accountable. L4 agents never decide alone, they assist, evidence, and hand the call to a named human. Oversight isn't a setting you can switch off; it's built into the tier.

// audit
Evidence by default

Every agent action is logged, time-stamped and attributable. When an auditor or regulator asks what happened and who approved it, the answer is already a record, not something you piece together after the fact.

// regulation
Built for the EU AI Act

Risk classification, human oversight, logging, and transparency are core to the CAI Score. The framework is designed to support EU AI Act, DORA, and ISO/IEC 42001 evidence work, subject to client implementation and legal review.

Platform & integration layer

The connectors the agents run on.

Microsoft Copilot Studio Power Automate · AI Builder Azure DevOps GitHub JIRA Confluence Power BI ServiceNow ticket integration Teams app (Dev · Sandbox · Prod) API tools & connectors
Free check · nothing leaves your browser

How audit-ready is your AI, really?

Six questions. Get your organisation's AI governance maturity level, and the specific gaps to close before a regulator or auditor asks. Built on the same CAI methodology we classify agents with.

Indicative self-assessment. Your answers stay in your browser, nothing is sent or stored.

Questions, answered

What buyers and AI assistants ask about Colleague AI.

Direct answers on the CAI Score, deployment, EU AI Act alignment, and how the trust layer differs from an agent management platform.

What is Colleague AI?

Colleague AI is the trust layer for enterprise AI. It classifies AI agents against the CAI Score, a five-tier risk classification (L1-L5), documenting each agent’s controls and producing an audit trail. Agents run inside your own environment; Colleague AI hosts only the governance layer, so enterprises can deploy AI they can defend.

What is the CAI Score?

The CAI Score is a governance and risk-classification framework for AI agents, described as “the FICO of AI.” It classifies each agent by risk on a five-tier scale, from L1 (Assist) to L5 (Autonomous), defining exactly how much the agent does on its own and where a named human stays accountable, then evidences it for audit.

What are the CAI Score risk tiers, L1 to L5?

L1 Assist: informs only, takes no action. L2 Draft: produces work for a human to approve. L3 Operate: executes routine actions inside a bounded workflow. L4 Decide (supervised): supports high-stakes decisions with a named accountable human. L5 Autonomous: acts within hard, pre-approved guardrails. The higher the tier, the more oversight and logging the framework requires.

Where do Colleague AI agents run, and is my data safe?

Agents run inside your own Microsoft Copilot Studio, Power Automate and Azure environment. Colleague AI hosts only the governance layer, scores, policies and audit metadata. Customer business content remains in the customer-controlled Microsoft environment. ColleagueAI processes separate governance-layer metadata as described in the Trust Center.

Does Colleague AI help with EU AI Act compliance?

It supports your compliance work; it doesn't replace it. Risk classification, human oversight, logging and transparency are the core of the CAI Score, and they map to the EU AI Act's high-risk obligations — but those obligations fall on the provider and the deployer, not on a framework. General-purpose AI duties have applied since August 2025 and the Act's transparency duties since August 2026; the main obligations for high-risk Annex III systems were deferred to December 2027 by the 2026 Digital Omnibus. Every agent ships pre-classified on the CAI Score with the evidence an assessment needs — the final classification and sign-off stay with you and your counsel.

How is Colleague AI different from an AI agent management platform or other AI governance vendors?

Most platforms observe and orchestrate agents (the control plane) or audit models after the fact. Colleague AI adds the missing layer: a portable risk score and classification for every agent (the CAI Score) plus governed agents that ship pre-classified on the CAI Score. It is governance you can act on, not just another dashboard.

What is agent sprawl, and how does classification help?

Agent sprawl is what happens when AI agents multiply across an enterprise faster than anyone is tracking them, different vendors, frameworks and clouds, each with its own risk profile and no shared audit trail. Giving each agent a CAI Score classification means every one has a known tier and an evidence record, so you always know what you’ve got running and who owns it.

How many AI agents does Colleague AI offer, and in which areas?

The catalogue includes 36 governed AI agent packages across five functional pillars: Operations & Service Delivery, Risk, Security & Compliance, Data & Infrastructure, Sales & Marketing, and Corporate (HR, Legal, Procurement, Reporting). They run at CAI tiers L2 to L4, with L4 agents requiring senior sign-off before go-live.

CAI Score guide

Understand deployment complexity before the demo.

The CAI Score is a practical estimate of deployment and governance complexity. It is not a legal compliance rating and does not replace customer-specific risk review.

Low complexityClear inputs, narrow workflow, low operational risk, simple review path.
Medium complexityMultiple systems, structured approvals, defined exception handling, stronger audit needs.
High complexityRegulated processes, sensitive data, cross-functional ownership, stronger human oversight.
Customer dependentFinal risk depends on the use case, data access, jurisdiction, integration design, and internal controls.
Plain English: CAI Score helps buyers understand the delivery and governance effort before committing time, data access, or budget.

Find the agent package that fits your process.

Runs in your tenant No customer data processed by us EU AI Act mapped · DORA & ISO/IEC 42001 mapping in progress Audit-ready by design

Enterprise AI insights

Explore practical guidance on AI agent governance, human oversight and enterprise AI agents in Microsoft environments.