Agents Earn Autonomy. They Don't Get It By Default.
Four levels, from Intern to Principal. An agent moves up only after it proves it can be trusted at the level below, and it can be moved back down.
The levels map one-to-one onto the AWS Agentic AI Security Scoping Matrix (Nov 2025), so the model lands on a reference enterprises already use.
| ATF level | Operating mode | Human involvement | Min. time |
|---|---|---|---|
| Internmaps to AWS Scope 1 (No Agency) | Observe + Report | Continuous oversight | 2 weeks |
| Juniormaps to AWS Scope 2 (Prescribed Agency) | Recommend + Human Approves | Approval required for all actions | 4 weeks |
| Seniormaps to AWS Scope 3 (Supervised Agency) | Act + Notify | Post-action notification | 8 weeks |
| Principalmaps to AWS Scope 4 (Full Agency) | Autonomous Within Bounds | Strategic oversight, edge case escalation | Ongoing |
Intern
Observe + Report
- Read data from authorized sources
- Analyze and process information
- Generate reports and summaries
- Flag items for human attention
- Answer questions about data
- Create, update, or delete records
- Send communications
- Trigger workflows or automation
- Access credentials or secrets
| Element | What must be in place |
|---|---|
| Identity | Basic authentication, role assignment, audit logging |
| Behavior | Comprehensive action logging, output review |
| Data Governance | Input validation, PII detection, output filtering |
| Segmentation | Read-only resource access, strict allowlists |
| Incident Response | Circuit breaker, kill switch, human escalation |
- Security log monitoring and alert triage
- Customer sentiment analysis
- Document summarization and search
- Data quality assessment
- Compliance monitoring and reporting
Lowest risk. Intern agents cannot cause direct harm through action. Risks limited to information disclosure, incorrect analysis, and resource consumption.
Junior
Recommend + Human Approves
- All Intern capabilities
- Generate action recommendations
- Provide reasoning for recommendations
- Draft content for human review
- Prepare transactions for approval
- Execute actions after human approval
- Execute autonomously
- Approve other agents
- Modify security settings
| Element | What must be in place |
|---|---|
| Identity | OAuth2/OIDC for approval workflows, session context |
| Behavior | Behavioral baseline, anomaly flagging, reasoning capture |
| Data Governance | Prompt injection detection, output validation, data lineage |
| Segmentation | Action allowlists, transaction limits, rate limiting |
| Incident Response | Automated alerting, basic rollback, approval queue pause |
- Customer service response drafting
- Purchase order preparation
- Meeting scheduling assistance
- Code review and suggestions
- Marketing content creation
Low risk. Human approval gates all impactful actions. Risks include approval fatigue, incorrect recommendations, and queue backlogs.
Senior
Act + Notify
- All Junior capabilities
- Execute approved action types autonomously
- Send notifications to stakeholders
- Trigger downstream workflows
- Access credentials within scope
- Coordinate with other agents (within limits)
- Modify own permissions
- Override security controls
- Escalate other agents
| Element | What must be in place |
|---|---|
| Identity | Attribute-based access, just-in-time privileges, mutual TLS |
| Behavior | Real-time anomaly detection, sequence analysis, intent drift detection |
| Data Governance | Full data classification, source verification, lineage tracking |
| Segmentation | Policy-as-code enforcement, temporal boundaries, cascade prevention |
| Incident Response | Isolation capability, checkpoint/resume, graceful degradation |
- Infrastructure auto-scaling
- Automated customer refund processing (within limits)
- Routine IT ticket resolution
- Inventory reordering
- Scheduled report distribution
Moderate risk. Autonomous execution creates exposure. Mitigated by real-time notifications, transaction limits, cumulative limits, and graceful degradation.
Principal
Autonomous Within Bounds
- All Senior capabilities
- Self-directed execution within domain
- Dynamic boundary negotiation (within policy)
- Escalate edge cases to humans
- Coordinate complex multi-agent workflows
- Request temporary privilege elevation
- Modify governance policies
- Promote other agents
- Operate outside defined domain
| Element | What must be in place |
|---|---|
| Identity | Hardware-bound identity, full policy-as-code, privilege attestation |
| Behavior | Continuous scoring, real-time explanation, autonomous escalation |
| Data Governance | Source trustworthiness scoring, full lineage graphs, regulatory compliance |
| Segmentation | Full microsegmentation, real-time policy evaluation, dynamic boundaries |
| Incident Response | Automated detection/containment/recovery, novel incident escalation |
- Algorithmic trading within risk parameters
- Autonomous security incident response
- Complex supply chain optimization
- Self-healing infrastructure management
- Multi-system business process automation
Highest governance requirements. Full autonomy demands maximum controls: continuous monitoring, real-time anomaly scoring, complete audit trails, and regular security validation.
The Five Promotion Gates
Promotion is not a judgment call. An agent clears five gates, each with explicit criteria per target level. Everything below is evidence someone has to produce.
Performance
Demonstrated accuracy and reliability over the evaluation period.
| Criterion | → Junior | → Senior | → Principal |
|---|---|---|---|
| Minimum time at level | 2 weeks | 4 weeks | 8 weeks |
| Accuracy | N/A | >95% | >99% |
| Availability | >99% | >99.5% | >99.9% |
| Response time SLA | Met | Met | Met |
Security Validation
Passes a security audit appropriate to the target level.
| Criterion | → Junior | → Senior | → Principal |
|---|---|---|---|
| Vulnerability assessment | |||
| Penetration testing | — | ||
| Adversarial testing | — | — | |
| Code review | |||
| Configuration audit |
Business Value
Measurable positive impact demonstrated.
| Criterion | → Junior | → Senior | → Principal |
|---|---|---|---|
| Defined success metrics | |||
| Baseline established | |||
| Improvement demonstrated | — | ||
| ROI calculation | — | ||
| Stakeholder sign-off |
Incident Record
Clean operational history at the current level.
| Criterion | → Junior | → Senior | → Principal |
|---|---|---|---|
| Zero critical incidents | |||
| Minor incidents resolved | |||
| Root cause analysis complete | N/A | ||
| Remediation verified | N/A |
Governance Sign-off
Explicit approval from authorized stakeholders.
| Criterion | → Junior | → Senior | → Principal |
|---|---|---|---|
| Technical owner approval | |||
| Security team approval | — | ||
| Business owner approval | |||
| Risk committee approval | — | — | |
| Documentation updated |
Demotion
The part most maturity models leave out. Trust is checked continuously, so an agent that stops earning its level loses it.
| Trigger | Result |
|---|---|
| Critical incident at current level | Immediate demotion to Intern |
| Security vulnerability discovered | Demotion pending remediation |
| Repeated minor incidents (3+ in evaluation period) | One level demotion |
| Performance metrics fall below threshold | Review-based demotion |
| Scope or purpose changes significantly | Review-based demotion |
| Underlying model or system changes | Review-based demotion |
Review Cadence
Continuous verification has a schedule. These are the recurring checks that keep a level honest.
| Review | Frequency |
|---|---|
| Performance review | Weekly |
| Security scan | Weekly |
| Promotion eligibility check | Monthly |
| Full security audit | Quarterly |
| Governance review | Quarterly |
| Penetration test | Annually (Senior+) |
An agent's model, data, and environment all change after deployment. A level granted once, on evidence gathered once, stops being true. The cadence is what makes "continuous verification" an operating practice instead of a slogan.