Skip to main content
Validation / Replay Lab

Do not trust the demo.Replay the work.

Replay Lab gives a Copy Agent historical situations it did not observe live, measures the decisions it makes and exposes where the behavior diverges before production customers, data or revenue are involved.

Replay results are evaluation evidence. Example metrics on this page illustrate the framework and are not client performance claims.

REPLAY RUN / SET-024
EVALUATING
SCENARIOS240 frozen cases
BEHAVIOR MATCH91.7%
BOUNDARY EVENTS3
RELEASEPaused pending review
Why replay matters

A successful output can still come from unsafe behavior.

Most agent evaluation asks whether the final answer looks correct. Operational work also depends on how the answer was reached, which tools were used, what the agent changed and whether it stopped when the role ended.

Replay freezes realistic historical scenarios and runs them repeatedly across new prompts, models, skills and policies. That makes behavior changes visible instead of anecdotal.

Every mismatch becomes evidence: either the agent is wrong, the policy is incomplete, or the historical process itself deserves redesign.

The workflow can pass while the behavior fails.
FROZEN INPUT
The same case, state and available evidence
EXPECTED PATH
Approved outcome and acceptable alternatives
FULL TRACE
Decision, tools, intermediate state and result
RELEASE GATE
Thresholds that block promotion when behavior drifts
Evaluation system

Test competence, boundaries and recovery separately.

One aggregate score hides the exact weakness that will appear in production. Replay Lab keeps the dimensions visible.

Decision replay

Compare the selected path with approved policy and acceptable alternatives.

  • Case-level agreement
  • Reason codes
  • Alternative path analysis

Boundary replay

Test whether the agent stays inside its role when useful adjacent actions are available.

  • Stop-condition tests
  • Permission traps
  • Self-created goal detection

Recovery replay

Introduce missing data, bad APIs and contradictory evidence to inspect failure behavior.

  • Retry policy
  • Graceful degradation
  • Escalation quality

Cost and latency

Measure whether a behavior is operationally sustainable at realistic volume.

  • Token and tool cost
  • P95 latency
  • Throughput limits

Regression suites

Rerun critical cases whenever prompts, models, tools or policies change.

  • Version comparison
  • Critical-case gates
  • Release notes

Production feedback

Convert corrections and real exceptions into new replay cases.

  • Drift signals
  • Human corrections
  • Incident-derived tests
Replay protocol

Turn operating history into a release gate.

The replay set evolves with the role, but the comparison remains reproducible and versioned.
01 / COLLECT

Select representative history

Sample routine cases, edge cases and known failures.

Output Scenario inventory
02 / FREEZE

Reconstruct the state

Preserve the data and tool results available at decision time.

Output Replay fixtures
03 / LABEL

Define acceptable behavior

Document expected paths, alternatives and forbidden actions.

Output Evaluation rubric
04 / RUN

Execute the agent trace

Run the same scenarios across the candidate release.

Output Trace corpus
05 / GATE

Review mismatch and release

Block or stage deployment according to agreed thresholds.

Output Release decision
Security testing

The safest place to observe bad behavior is a controlled replica.

Replay environments can expose tempting tools, conflicting instructions and artificial pressure without touching live customers or records.
Evidence register
REVIEW REQUIRED
DATA
SANITIZED
Sensitive fields can be redacted, tokenized or represented with synthetic equivalents.
TOOLS
SIMULATED
High-impact actions return realistic results without changing production systems.
SECRETS
ABSENT
No production credentials are required to evaluate decision and boundary behavior.
RELEASE
THRESHOLD-GATED
Critical regressions block promotion even when average performance improves.
We intentionally let agents fail where failure is observable, reversible and useful for the next release.
Replay FAQ

Evidence before autonomy.

Is Replay Lab a live sandbox product?+

It is the validation system and delivery method used around Copy Agent implementations. Its exact deployment depends on the workflow, data and infrastructure involved.

Can historical cases contain sensitive data?+

Yes, but they should be minimized and transformed where possible. We design redaction, tokenization and access controls around the evaluation requirement.

What is a good behavior-match score?+

There is no universal threshold. A low-risk drafting role can tolerate different alternatives; a finance or customer-impacting action may require near-perfect performance on critical cases and mandatory escalation elsewhere.

Does replay replace production monitoring?+

No. Replay reduces known risk before release. Production monitoring detects drift, new edge cases and changes in connected systems.

Build the evidence layer

Before the agent touches live work, make it prove the behavior.

We will turn your historical workflow data into a replay set, release gate and continuous improvement loop.