How does an AI system know when its work is actually complete?
An emerging research direction on executable acceptance criteria, automated validation and repair, and consistent assumptions across durable agent workflows.
The problem
AI systems can produce plausible outputs without demonstrating that requirements are covered, tests pass, or assumptions remain consistent across tools and retries.
The approach
Study self-testing, adversarial test generation, invariant checking, requirement coverage, checkpoint semantics, provenance, compatibility, and explicit stopping criteria.
An emerging research direction on executable acceptance criteria, automated validation and repair, and consistent assumptions across durable agent workflows.
Main contributions
- Research framing and system design
- Methods, implementation, and empirical evaluation
- Open research artifacts and scholarly dissemination