Add Memrail to an existing application one consequential decision at a time. Begin with a read-only topology audit, compare policies in an isolated evaluation project, and enable the new execution path only after its authorization and fallback behavior are tested.
Start with discovery, not extraction#
Run a read-only decision topology audit before changing runtime behavior. The report should identify:
- every consequential point where context changes what the system does;
- every model output that is treated as permission rather than evidence;
- state, tags, and events available at each point;
- current tool executors, routes, human approvals, and failure paths;
- project boundaries and cross-project dependencies;
- branches where no explicit default or denial behavior exists.
This separates where a decision happens from how it is currently expressed. Several scattered if statements may belong to one named decision point; one agent loop may need several control points because it proposes tools, data access, and user-facing responses at different stages.
1. Detect the integration surface#
Check the project language and whether Memrail is already present:
rg -n 'memrail|@memrail/sdk|AsyncAMIClient|AMIClient|\.decide\(' \
package.json pyproject.toml requirements*.txt src app lib 2>/dev/null
Use the Python SDK if the Python package is present and the TypeScript SDK for @memrail/sdk. A mixed repository can use both, but project ownership and event naming must remain coherent.
2. Inventory hidden authority#
Search patterns are clues, not proof. Review each result in context:
rg -n --glob '*.{py,ts,tsx,js,jsx}' \
'if .*tier|if .*priority|score.*[><=]|threshold|feature.?flag|variant|ab.?test'
rg -n --glob '*.{py,ts,tsx,js,jsx}' \
'tool_choice|function_call|execute_tool|send_email|refund|approve|deny|escalat|route'
rg -n --glob '*.{py,ts,tsx,js,jsx}' \
'system_prompt|systemPrompt|instructions|guardrail|moderation|classifier|predict\('
Classify findings into three categories:
| Category | Example | Migration treatment |
|---|---|---|
| Evidence | an LLM classifies intent | expose as a constrained tag |
| Authority | a score above 0.8 issues a refund | move the threshold and action selection into an EMU |
| Execution | a refund client calls the payment API | keep as a registered, versioned tool executor |
The goal is not to remove inference. It is to prevent inference from silently acquiring authority over state changes.
3. Select one control seam#
Choose a decision that is consequential enough to matter and bounded enough to test. Good first seams include:
- whether an agent may call a side-effecting tool;
- whether a case is routed to a human;
- which response protocol applies to a regulated or high-value user;
- whether a workflow may advance to its next irreversible state.
Avoid beginning with a global middleware gate unless all downstream paths share one atom contract and one failure policy. A named hook close to the action is usually easier to reason about.
decision_point: billing.refund
current_authority:
file: src/agents/billing.py
behavior: model confidence above 0.82 calls refund tool
proposed_boundary:
model_role: classify request and extract bounded facts
memrail_role: choose allow, require_human, or deny route
executor_role: perform refund only for an authorized tool_call
4. Compare in an isolated evaluation project#
Add an evaluation-only call on your AsyncAMIClient without changing the existing branch. Define a shadow EMU that expresses the current rule and compare its trace with actual application behavior. Use a separate project to keep migration fixtures and observations distinct.
decision = await client.decide(
decision_point="billing.refund",
context=[
state("customer.id", customer.id),
state("refund.amount", amount),
state("refund.currency", currency),
tag("refund_reason", classified_reason, source="ml"),
],
project="billing-evaluation",
options=InvokeOptions(dry_run=True),
trace=TraceOptions(enable=True),
)
# Existing behavior continues during the comparison period.
result = await legacy_refund_branch(...)
Bind candidate policies to billing.refund. Validate classifier enums at runtime, preserve source labels, and keep authoritative facts application-owned.
Keep the existing branch as the sole executor during comparison. Inspect trace reasons, score, and suppressed_by under the evaluation contract.
5. Prove trigger reachability#
For every EMU proposed at the seam, create an atom contract table:
| Trigger dependency | Source | Always present? | Type/domain | Failure behavior |
|---|---|---|---|---|
state.refund.amount |
request validator | yes | number ≥ 0 | reject malformed request |
state.customer.id |
authenticated session | yes | string | deny if absent |
tag.refund_reason |
bounded classifier | no | fixed enum | require human if absent |
event.customer.received.refund |
payment success event | after rollout | configured retention (default 90 days) | duplicate guard unavailable before complete instrumentation |
Positive predicates on absent state/tag atoms evaluate false; outer NOT can invert this, and OR can match another branch. Guard required IDs with EXISTS, including IDs interpolated into event filters. A negated event that is never emitted becomes true. Audit both positive and negative dependencies, retention, and producer coverage. Observed schema validation alone does not prove every caller supplies the required atoms.
6. Connect actions explicitly#
Build an action connectivity matrix before activation:
action_connectivity:
context_directive:
status: connected
handler: build_billing_system_prompt
decision_prompt:
status: connected
handler: refund_review_queue
tool_call.refund_payment:
status: connected
executor: payments.refund
version: 1.0.0
idempotency: refund_request.id
route.billing_manual_review:
status: missing
An EMU that references a missing executor is not complete. Verify the project, tool version, schema, and real runtime handler. The SDK executor enforces policy/lifecycle checks; custom dispatchers must preserve that metadata and apply equivalent checks. Neither dry-run nor advisory tool selections may become live execution merely because they were returned. Handler permissions and business duplicate protection remain application responsibilities.
7. Cut over one branch at a time#
Promote only after comparison data shows equivalence or an intentional difference. A safe sequence is:
- instrument context and events;
- evaluate the equivalent policy in an isolated shadow project;
- make the selected action observable but non-executing;
- route an explicit application-owned cohort to a reversible or human-reviewed branch (
canarystate alone does not sample traffic); - route all traffic through Memrail at that decision point;
- remove the legacy authority after the rollback window;
- keep the executor and domain validation in application code.
The new boundary should fail closed for dangerous actions and preserve a documented default for ordinary flow. Network failure must not silently hand authority back to the model.
Definition of done#
- Every consequential path reaches the named decision point or is explicitly out of scope.
- The atom contract covers every trigger dependency and interpolation value.
- All model-derived tags use fixed taxonomies with explicit unknown/failure handling.
- Evaluation-only mode and application authorization are enforced before every dispatch, independently of SDK selection.
- Every selected action type has an application handler.
- Every tool is registered, versioned, and in the same project as its EMUs.
- Material outcomes emit events with entity identifiers needed by
WHEREclauses. - Shadow traces have been compared against real behavior.
- Rollback restores a safe known behavior, not implicit model authority.
- Production EMUs and
.emu.lock.jsonlare committed and reviewed.