Skip to main content

Agent Harness

The Harness is the core of cascadeflow’s runtime intelligence. It wraps agent execution and makes a decision at every step — should this model call proceed, be switched, or be stopped?

What the Harness Does

At every LLM call or tool execution inside an agent loop, the Harness:
  1. Checks hard constraints — budget remaining, compliance allowlist, tool call cap, latency limit, energy limit
  2. Scores soft dimensions — quality, cost, latency, energy weighted by KPI priorities
  3. Decides an actionallow, switch_model, deny_tool, or stop
  4. Records a trace — action, reason, model, step, cost, budget state
In observe mode, decisions are recorded but not enforced. In enforce mode, they shape execution in real time.

HarnessConfig — The Full Control Surface

All Harness behavior is configured through a single dataclass:

The Three-Tier API

cascadeflow offers three levels of control — use the one that fits your needs:

Tier 1: Global Init (Zero-Change)

Best for: first rollout, measuring baseline costs, auditing compliance.

Tier 2: Scoped Run (Block-Level Control)

Best for: per-request budgets, scoped policy, session-level metrics.

Tier 3: Agent Decorator (Per-Agent Policy)

Best for: multi-agent systems where each agent has different constraints.

Decision Priority

When the Harness evaluates a step, it follows a strict priority order: Hard constraints (budget, compliance) always take priority over soft scoring (KPI weights).

Six Dimensions at a Glance

Observe vs Enforce

Start with observe to validate your policies against real traffic. Switch to enforce when you are confident the rules are correct.
Run this example: examples/enforcement/basic_enforcement.py | API reference: HarnessConfig

Next Step

See how the Harness operates inside multi-step agent loops. Understand the Agent Loop →