The Execution Layer for
Enterprise AI Agents.
WorkAgent OS is not a chatbot. It is a specified multi-tenant execution architecture where human knowledge workers and AI agents collaborate through live company context, policy-controlled tools, verifiable execution, and human-in-the-loop approvals.
The architecture is designed to give AI agents identity, context, scoped tools, policy controls, approvals, verification, and auditability — so they can safely act inside real business workflows.
Public specification v1.1 — production implementation and verification evidence have not been published with this website.
* The model is not the product. Models are interchangeable reasoning modules inside the Agent Runtime.
How WorkAgent OS Executes Work
From authenticated user identity to verified external action, every operation follows a specified 8-layer enterprise lifecycle:
Identity & Context
Establishes identity chain (Tenant → User → Agent) and curates minimal necessary context.
Agent Runtime
Interchangeable reasoning engine orchestrating intent, memory, and DAG steps.
Planning DAG
Compiles dependency graph; runs independent retrieval sub-tasks in parallel.
Tool Selection
Selects tools matching strictly defined JSON Schemas via the MCP-native gateway.
Risk & Policy
Dynamic Risk Formula (Impact × Irreversibility × Externality × Sensitivity × Uncertainty × 100).
Action Execution
Canonical payload hashing with unique idempotency keys to eliminate side-effects.
Verification
Active external state read-back comparing actual tool outcome against expected schema.
Audit Ledger
Emits tamper-evident audit event linking actor, canonical action hash, and result.
Built for Every Enterprise Stakeholder
WorkAgent OS aligns the productivity needs of employees with the rigorous governance mandates of IT and Security teams.
Personal AI Chief of Staff & Productivity
"My AI Chief of Staff acts like a senior executive assistant—reading my schedule, finding the right Drive slides, and preparing draft tasks before I even ask."
From Manual Knowledge Work to Verified Execution
How enterprise workflows are designed to transform when autonomous agents operate within policy boundaries, verifiable tool access, and action integrity hashing. Platform capabilities below are designed targets, not measured benchmarks. Read the specification.
| Operational Dimension | Traditional Enterprise Workflow | WorkAgent OS Execution Platform |
|---|---|---|
01.Context Assembly | Analysts manually cross-reference Gmail threads, Drive files, Slack conversations, and Jira tickets — a slow, repetitive process prone to stale data. | Designed: automated 9-stage context resolution plus a separate retrieval/ranking pipeline (structured fetch, semantic search, relevance, source authority, recency, dedup, token budget, citations). |
02.Tool Writes & Mutations | Copy-paste between tools, forgotten CRM stages, mistyped ticket updates, and orphaned follow-ups prone to human fatigue. | Designed: idempotent execution DAG, pre-validated against tool JSON schemas via an MCP Gateway so credentials stay out of model context. |
03.Risk Scoring & Approvals | Informal Slack DMs or unreviewed auto-execution scripts; zero quantitative risk assessment before high-impact changes. | Designed: normalized Risk Formula (0–100). Actions scoring > 50 are intended to require hash-bound human approval before execution can proceed. |
04.Failure Recovery | Silent failures on 429/500 API responses; half-executed workflows left in corrupted states requiring manual engineering triage. | Designed: active state verification loop — re-reads external system state and triggers retry, compensating action, reconciliation, or escalation depending on connector capability. |
05.Forensic Audit Ledger | Fragmented logs siloed across 5 separate vendor portals; zero tamper-evident link connecting original user intent to side-effects. | Designed: tamper-evident audit ledger, sequentially chained with SHA-256 hashes binding tenant, actor, policy version, and canonical action payload. |
“Execution isn’t success until external state is verified.”
The verification principle behind the proposed read-back and recovery model.
“External content can provide information. It cannot grant authority.”
The trust principle behind the proposed authority hierarchy and untrusted-content firewall.
AI Chief of Staff — Preview Overview
Acts as a personal autonomous chief of staff for employees and executives. Reads calendar schedules, searches email threads, synthesizes Drive proposals, catches delayed deliverables, and drafts action plans.
Production implementation and verification evidence have not been published with this website. The Chief of Staff preview is an early demonstration of intended behavior — it is not a live production deployment.
- Daily executive priority briefing and blocker detection
- Meeting prep with automatic CRM, Gmail, and Drive synthesis
- Action item extraction and Jira/Linear task generation
- Overdue deliverable tracking across team tools
The 6-Layer Platform Architecture
WorkAgent OS decouples interchangeable reasoning models from business rules, persistence, and tool integrations. The 6-Layer Platform Architecture provides the multi-tenant infrastructure host, while the 8-Layer Execution Model governs each agent action lifecycle.
Presentation & Workspace Layer
Modern Next.js 16 web application, role-based Employee Workspace, Admin Console, and Human-in-the-Loop review portals.
API Gateway & Security Boundary
Enterprise gateway (SPEC) designed for authentication, request-signature validation where applicable, tenant isolation, and rate limiting. Request-signature checks (e.g. JWT or webhook verification) are distinct from action payload hashes.
Agent Runtime & Execution Engine
The core reasoning engine. Features modular LLM adapters (Claude, GPT, DeepSeek), 9-step Context Engine, Dynamic Risk Engine, and State Verification.
MCP-Compatible Tool Gateway
Bi-directional connector gateway providing tenant-scoped, permission-controlled tool access to enterprise SaaS and internal databases.
Enterprise Multi-Tenant Storage Layer
Strictly segregated multi-tenant persistence layer combining relational data, vector embeddings, and tamper-evident audit logging.
Observability, Cost & Evaluation Layer
Full OpenTelemetry-compliant execution tracing, real-time cost accounting, and automated regression evaluation suites.
Agent Runtime & Execution Engine
The core reasoning engine. Features modular LLM adapters (Claude, GPT, DeepSeek), 9-step Context Engine, Dynamic Risk Engine, and State Verification.
Core Modules & Protocols:
9-Stage Context Resolution
Context resolution stages: Identity, Tenant, User, Task, Conversation, Org, Memory, Tool, Policy. Retrieval and ranking (structured fetch, semantic search, relevance, source authority, recency, dedup, budget, citations) run as a separate pipeline.
Planning & Tool Router
Decomposes requests into step-by-step DAGs, selects verified tools, validates JSON schemas.
Policy & Risk Engine
Normalized formula: Impact × Irreversibility × Externality × Sensitivity × Uncertainty × 100. 0-100 tiered action gates.
State Verification & Recovery
Re-reads external state post-execution, verifies outcome schema against expected entity, and triggers compensation, retry, reconciliation, or escalation depending on connector capability.
The WorkAgent OS Agent Fleet
Led by the AI Chief of Staff as the proposed product wedge (PREVIEW), backed by specialized workforce agents defined in the specification with strict tool contracts and scoped permissions.
AI Chief of Staff
PREVIEWProposed WedgeExecutive Co-pilot & Orchestrator
Proactive daily prioritization, cross-tool context synthesis, and executive meeting prep.
"Prepare me for today's 2 PM Acme executive briefing and flag any open Jira blockers."
Key Responsibilities:
- Daily executive priority briefing and blocker detection
- Meeting prep with automatic CRM, Gmail, and Drive synthesis
- Action item extraction and Jira/Linear task generation
- Overdue deliverable tracking across team tools
- Weekly automated executive progress digests
Execution Reasoning Trace (DAG Steps):
Enterprise Sales Agent
SPECPipeline Acceleration & Deal Intelligence
Autonomous account research, CRM hygiene, and meeting brief generation.
"Summarize Acme Corp's deal trajectory, recent objection trends, and recommend next pricing tier."
Customer Success Agent
SPECRetention & Proactive Health Monitoring
Monitors client health scores, surfaces churn risks, and orchestrates quarterly business reviews.
"Run a sentiment and health analysis for our top 10 enterprise accounts renewing this quarter."
Project & Sprint Agent
SPECDelivery Orchestration & Blocker Resolution
Cross-functional sprint coordination, dependency graph analysis, and velocity optimization.
"Identify which PRs are blocking the upcoming v2.4 release and ping relevant reviewers."
Autonomous Meeting Agent
PREVIEWLive Transcription & Consensus Capturing
Real-time meeting synthesis, commitment tracking, and instant bi-directional tool sync.
"Extract decisions from today's Product Architecture Sync and create tickets for owners."
Deep Research Agent
SPECMarket Intelligence & Multi-Source Synthesis
Exhaustive multi-source investigation with tamper-evident source citations and fact-checking.
"Conduct a comparative teardown of modern Enterprise MCP gateway architectures."
Developer & Codebase Agent
SPECCode Intelligence & CI/CD Debugging
Autonomous codebase navigation, PR triage, test reproduction, and safe refactoring.
"Diagnose intermittent timeout in AgentRun webhook test and generate fix PR."
Finance & Operations Agent
SPECCompliance & Budget Governance
Expense reconciliation, budget tracking, vendor invoice verification, and risk auditing.
"Audit enterprise API spend for last month and flag any department exceeding allocated quota."
- No production side-effects — nothing executes against real systems.
- Identities, tool calls, hashes, costs, and timings are examples.
- Clicking approve performs no real authentication or hash verification.
Target semantics are documented in the Trust Center approval lifecycle.
End-to-End Execution & Approval Trace
Witness the exact execution sequence defined in Section 23 of the master spec: "Prepare me for today's Acme executive meeting" including context assembly, risk evaluation, and hash-bound human sign-off.
AI Chief of Staff (Tenant: enterprise_corp_89)
Simulated authentication of session user ([email protected]) under Tenant ID tenant_89f2a. Example identity delegation chain: Tenant → User → Agent.
{"tenant_id": "tenant_89f2a", "user_role": "VP_PRODUCT", "identity_chain": "tenant_89f2a:user_alex:agent_chief_of_staff", "status": "AUTHORIZED (simulated)"}The Dynamic Risk Scoring Engine
The specification defines a normalized quantitative score that is evaluated before execution. This is an illustrative scoring model — policies, tenant configuration, and scope checks can still deny low-scoring actions:
Execution paused. Canonical JSON SHA-256 action hash locked until authorized user sign-off.
The WorkAgent OS Data Model
Formally defined in Section 20 & 21 of the specification. Built for PostgreSQL with Row-Level Security (RLS) to enforce multi-tenant boundaries at the database kernel level.
AgentRun
Single end-to-end execution lifecycle instance capturing telemetry and token costs.
Columns & Types:
| Column | Type | Description |
|---|---|---|
| id | UUID | Primary key |
| conversation_id | UUID | Parent conversation |
| status | ENUM | PENDING, RUNNING, APPROVAL_WAIT, COMPLETED, FAILED |
| started_at | TIMESTAMPTZ | Execution start |
| completed_at | TIMESTAMPTZ | Execution completion |
| cost | NUMERIC(10,5) | Total token cost in USD |
Foreign Key Relationships (Section 21 ERD):
REST & MCP API Surface
Section 22 of the specification defines proposed endpoints for agent runs, hash-bound approvals, semantic memory queries, and audit logs. All endpoints below are SPEC — responses are SIMULATION examples, not live latency or availability guarantees. Copying a request payload does not imply an implemented API.
Authorization: Bearer <tenant_scoped_jwt>
X-Tenant-ID: tenant_89f2a
Proposed endpoint: initialize an autonomous execution run for an agent with tenant context and prompt payload.
{
"conversation_id": "conv_9921",
"prompt": "Prepare me for today's Acme meeting and summarize blockers.",
"context_overrides": {
"time_window_hours": 48
},
"max_budget_usd": 0.50
}{
"run_id": "run_01j98ab7",
"agent_id": "chief-of-staff",
"status": "RUNNING",
"estimated_cost_usd": 0.084,
"steps_planned": 5,
"created_at": "2026-10-08T18:30:00Z"
}Definition of Done — Readiness Targets
Section 30 of the master specification defines the readiness targets an agent must meet before any production release. These are engineering targets in the specification — no item below has been verified in a published production deployment. Production implementation and verification evidence have not been published with this website.
Monorepo setup, Auth, Tenant Resolver, PostgreSQL schema, AgentRun tables, Claude adapter, baseline UI.
Google Calendar, Gmail, Drive connectors, Tool registry, Context Builder ranking, Chief of Staff MVP.
Slack and Jira/Linear MCPs, ABAC permission engine, approval UI, idempotency keys, audit events.
OpenTelemetry tracing, token cost accounting, golden evaluation suite, prompt injection tests, production demo.
Proposed Connector Catalog
Every connector referenced by the agent fleet is listed here with its public status. These are proposed connectors — design examples only; no production connector registry has been published and none are silently active.
Honest Status, Data Handling & Security Posture
Status legend for every claim on this site, identity and approval lifecycles, residency and retention matrices, model provider handling, threat model, and responsible disclosure — all labeled LIVE, PREVIEW, SPEC, or ROADMAP.
Ready to Deploy Autonomous Agents with Enterprise Confidence?
The WorkAgent OS specification unites AI agents, enterprise SaaS tools, and human oversight into one policy-bounded, verifiable execution architecture.