LEARN FROM OUTCOMES • AMI · DSL v2

Recursive self-improvement with Hindsight

Improve agent steering and control through a reviewed feedback loop. Use Memrail Hindsight to analyze outcomes, propose policy changes, and evaluate the next version.

View raw Markdown

Memrail Hindsight enables recursive agent improvement by turning recorded behavior into proposals for better policy. Analyze what happened, review a proposed change, deploy a tested policy version, and observe its effects. Those new observations become evidence for the next cycle.

The thing that improves is the agent's explicit steering and control policy—not its model weights. Hindsight supplies retrospective analysis and proposals; your application and release process supply outcome measurement, authorization, and controlled deployment. Improvement is a hypothesis to test, not a guarantee of running the loop.

What makes the loop recursive?#

An EMU (Executable Memory Unit) is a versioned condition/action policy. A decision trace records an evaluation; an event records an occurrence supplied by your application. Hindsight examines these artifacts and proposes changes to the policy that will shape subsequent behavior.

text
Policy version → decisions → authorized actions → recorded outcomes
      ↑                                                ↓
Reviewed release ← tests and review ← proposal ← Hindsight analysis

This is more than saving a reflection in a prompt. A useful lesson becomes an inspectable rule with named inputs, a condition, and a prescribed action. The next cycle examines behavior under that changed rule. Runtime trigger evaluation remains deterministic for its complete inputs; Hindsight's LLM-assisted interpretation is a separate, probabilistic step.

Consider a support agent that repeatedly sends billing cases to a general queue. This illustrative sequence shows how the loop can improve routing without expanding the agent's permissions:

Cycle Evidence to examine Bounded change to test
Establish a baseline Routing traces plus application events show billing cases being transferred twice Propose a billing-route EMU using a validated intent tag and trusted queue-availability facts
Evaluate the new policy New outcomes show fewer transfers, but some cases reach a queue that cannot serve their language Refine the routing condition to require a supported language; retain the general-queue fallback
Check the refinement Compare later cases with the baseline and held-out cases Keep, revise, or revert the change based on correct routing and resolution—not merely how often the new rule matches

These are example hypotheses, not promised Hindsight findings. A reviewer must distinguish a missing policy from a bad classifier, missing input, or disconnected executor. Sometimes the right improvement is repairing an input producer, not adding another rule.

What does Hindsight analyze?#

Analysis runs asynchronously at workspace scope, with three selectable pipelines:

Pipeline Evidence and questions
event_analysis Event co-occurrence and coverage gaps: which patterns recur, and which lack corresponding policy? Deterministic statistics help ground the LLM's interpretation.
emu_analysis Existing policies: where are rules redundant, stale, or potentially conflicting?
trace_analysis Decision traces and available policy context: which conditions fail, which rules are never selected, and where might behavior need a new or revised policy?

An insight includes a pattern description, confidence scores, and available evidence references. Trace insights can include sample trace IDs; other insights can reference events and EMUs. Proposals link back to insights and can propose creating, merging, or archiving EMUs. Inspect proposed_emu for create/merge candidates and the target references for archive candidates.

These are not all equally strong forms of evidence. Co-occurrence is not causation; a rarely selected rule may be an important safeguard. Confidence and expected utility are estimates, not proof that a change improves business outcomes.

Record evidence the next cycle can use#

Enable decision tracing and emit events after material actions. Capture the actual result: executed, failed, declined, or still uncertain. A selected action in a decision trace does not establish that the action ran or helped the user. See Python tracing and events or TypeScript tracing and events.

Keep request/correlation identifiers and policy versions in your application records so reviewers can connect decisions to later outcomes. For routing, measure transfers and resolution; for refunds, measure duplicate prevention and reconciliation as well as successful payments. These outcome definitions and joins belong to your application; Hindsight does not infer a trustworthy reward function for you.

Send only data you are authorized to use for LLM-assisted analysis. Redact unnecessary personal data and secrets before ingestion, and treat user-authored event text as evidence, not instructions to the reviewing agent. Analysis cannot recover expired history: raw decision traces have a default seven-day retention; event retention is organization-configurable. See events and traces.

Start an analysis through the API#

Use a workspace with relevant recorded activity and a deployment with Hindsight processing enabled. Set AMI_API_KEY, AMI_ORG, AMI_TEAM, and AMI_WORKSPACE for that workspace; set AMI_BASE_URL to your API origin, such as https://api.memrail.com (without /v1). Keep the key in your environment, not in a prompt or repository.

This request queues analysis and can consume LLM tokens. It creates insights/proposals; it does not activate policies:

bash
curl --fail-with-body --silent --show-error \
  --request POST \
  "${AMI_BASE_URL}/v1/workspaces/${AMI_WORKSPACE}/hindsight/analyze" \
  --header "Authorization: AMI-Key ${AMI_API_KEY}" \
  --header "X-AMI-Org: ${AMI_ORG}" \
  --header "X-AMI-Team: ${AMI_TEAM}" \
  --header 'Content-Type: application/json' \
  --data '{"pipelines":["event_analysis","emu_analysis","trace_analysis"],"analysis_window":"P7D","force":false}'

The API returns HTTP 202 with job_id, initial status: "pending", and a status_url. Poll that URL on the same API origin using the same authentication and organization/team headers. Jobs move through pending and processing to completed, failed, or cancelled. An already-active workspace job normally returns 409; wait for it instead of forcing overlapping work.

All paths below are relative to the API origin and require the same headers. Replace {workspace}, {job_id}, and {proposal_id} with the relevant values.

Method and path Purpose
GET /v1/workspaces/{workspace}/hindsight/jobs/{job_id} Poll progress and retrieve summary counts or an error
GET /v1/workspaces/{workspace}/hindsight/jobs Inspect full job results, including skipped pipelines and budget-limited completion
GET /v1/workspaces/{workspace}/hindsight/insights Read detected patterns and supporting evidence
GET /v1/workspaces/{workspace}/hindsight/proposals?state=PROPOSAL&source=hindsight Read candidates awaiting review
GET /v1/workspaces/{workspace}/hindsight/proposals/{proposal_id}/validation Read a stored validation report, if present; 404 can mean no report exists yet
POST /v1/workspaces/{workspace}/hindsight/proposals/{proposal_id}/validate Refresh the proposal's validation report without deploying or promoting it

List responses are paginated with limit, offset, and total. A completed job is not a coverage certificate: pipelines may be skipped or processing may stop at a budget limit. Inspect the full job's result.skipped_pipelines and result.budget_exceeded, not only its status or proposal count.

Turn a proposal into a reviewed policy change#

Generated candidates start in proposal state PROPOSAL. They are not active EMUs. Refresh and inspect the proposal's validation report, then validate the exact JSONL candidate through emu-plan --strict before release. The atom schema registry records observed input schemas; checks against it do not establish business correctness, authorization, or that every caller supplies those inputs. Revalidation updates feedback, not deployment approval.

For production, use Hindsight as a source of candidate changes for the JSONL workflow:

  1. Inspect the insight's evidence, the proposed rule, and the exact workspace/project and affected EMUs. Keep the insight/proposal IDs with the review.
  2. Translate the intended change into emus/emus.jsonl; do not copy the entire proposal envelope as a deployable EMU. Verify tool IDs/versions, trusted atom sources, policy fields, and literal-key idempotency for side effects.
  3. Run positive, negative, missing-input, competing-policy, and executor tests. For a merge or archive, demonstrate which existing behavior remains covered.
  4. Review emu-plan --strict, which validates the complete candidate before writes, and evaluate in staging with dispatch disabled. Release only with explicit approval, application-controlled rollout, outcome checks, and a rollback plan.

The proposal lifecycle API can also register EMUs and change their lifecycle. Treat those operations as policy writes, not harmless review labels; do not assume it enforces your staged approval process. Avoid mixing direct promotion with a Git-managed production policy. See lifecycle behavior for the limits of shadow and canary states.

Give a coding agent a bounded assignment:

text
Review this Hindsight insight and proposal against the current application.
Verify the evidence and input producers. State the expected outcome change.
Prepare the smallest policy-as-code diff and regression tests, including
competing policies and executor authorization. Preserve permissions,
approval requirements, duplicate protection, and fallback behavior.
Do not apply, promote, archive, enable dispatch, or deploy.
Flag any additional required change before making it.

The reviewable refund change shows what this contract looks like for a concrete policy edit.

How do later cycles avoid repeating the same work?#

Run subsequent analysis from your own workflow after enough new evidence arrives. Hindsight uses input watermarks and processed-record tracking to reduce repeated analysis, plus canonical pattern/proposal deduplication. Proposal generation skips patterns already proposed or rejected; optional semantic deduplication can reduce similar suggestions. Repeated runs are incremental, not guaranteed full replays of the requested window.

That makes iteration more manageable, but not exhaustive. Check coverage and retain a separate evaluation dataset. Rejection history helps avoid repeating a proposal; it is not model retraining or proof that the model learned a general constraint. Define how your team will reconsider a rejected hypothesis when materially new evidence arrives.

Can the improvement process itself improve?#

Hindsight uses versioned analysis prompts and records prompt/model provenance on insights and proposals. That gives you a second reviewable surface: compare analysis versions against a fixed dataset and assess evidence quality, duplicate suggestions, missed cases, and review cost.

Changing those prompts is a separate controlled change. Hindsight does not automatically rewrite its own prompts, objectives, permissions, or model weights. Keep approval authority and evaluation criteria outside the proposing agent's control.

Start with one decision point, one outcome measure, and one reviewed candidate. After release, record what actually changed and feed that evidence into the next analysis. The loop is recursive because each tested policy changes the behavior the next cycle studies—not because the system is allowed to approve its own conclusions.