DIAGNOSIS • AMI · DSL v2

Troubleshooting

Diagnose Memrail policies that do not fire, actions that cannot execute, validation failures, event mismatches, configuration errors, and unsafe fallbacks.

View raw Markdown

To diagnose a missing Memrail action, separate trigger matching, policy selection, and application execution. Inspect decision traces for matches and suppressions, and application logs for what actually ran. Keep application dispatch disabled during diagnostic evaluations.

An EMU does not fire#

Enable dry run and tracing at the call site, with application dispatch disabled. Pass the policy's named decision point:

python
from memrail.atoms import state
from memrail.models import InvokeOptions, TraceOptions

response = await client.decide(
    decision_point="support.triage",
    context=context,
    options=InvokeOptions(dry_run=True),
    trace=TraceOptions(enable=True),
)

Then check:

  1. Is the EMU in a state that is evaluated?
  2. Does the EMU's decision_point exactly match the named call?
  3. Is the call scoped to the expected organization, workspace, and project?
  4. Is every state and tag dependency present?
  5. Are the values and types capable of satisfying the operators?
  6. Do event names, attributes, and timestamps match?
  7. Was the match suppressed by cooldown, idempotency, or exclusion?

Run project validation:

bash
memrail emu-validate -w production -p support-agent

Missing atom#

Memrail DSL
// Trigger requires two inputs.
state.customer.tier == 'enterprise' AND tag.intent == 'refund'

If the decision call sends only the tier, this positive conjunction evaluates false. That does not mean every trigger with a missing input is false: NOT reverses a false predicate, and another OR branch may match. In particular, NOT state.user.banned is true when the key is missing. Use explicit presence and boolean contracts for authorization:

Memrail DSL
state.user.banned EXISTS AND state.user.banned == false

Conditional omission#

A builder may omit the exact value needed for the exceptional case:

python
# Broken for users who never logged in.
if user.last_login_at is not None:
    context.append(state("user.days_since_last_login", days))

Represent “never” explicitly with a separate boolean or documented sentinel. Do not assume missing is equivalent to zero, false, or infinity.

Wrong atom kind#

state("ticket.priority", "high") does not satisfy tag.priority == 'high'. Match the trigger namespace and source semantics; state keys also need at least two namespaced segments.

A negative event guard always passes#

Memrail DSL
NOT event.agent.sent.email IN 'P7D'

If no producer emits agent.sent.email, this condition is always true. Search emitters and compare exact name order:

bash
rg -n 'emit_event|ingest_event|emitEvent|agent.*sent.*email|email.*sent.*agent' src

Emit the canonical event after the email service confirms success. Do not “fix” the trigger by removing the duplicate guard.

Also check event retention, ingestion lag, scope, and missing filter identifiers. Ninety days is the default organization event-retention setting, not a fixed maximum. A negative query over expired or never-collected history cannot establish that the action never happened.

An event matches the wrong entity#

Unscoped event queries may see another entity’s event. Add an attribute at emission and filter it with a template literal:

Memrail DSL
state.customer.id EXISTS
AND event.agent.sent.email
  WHERE customer_id == '{{customer.id}}'
  IN 'P7D'

WHERE customer_id == state.customer.id is invalid. The event attributes must actually contain customer_id, with the same string value as the resolved template. Missing placeholders stay literal and normally match nothing; a surrounding NOT would then evaluate true. The explicit EXISTS guard prevents missing context from becoming permission.

An action is selected but nothing happens#

First distinguish “selected” from “executed.” A decision trace proves policy selection, not external side-effect completion.

bash
memrail action-connectivity -w production -p support-agent
memrail tool-get-schema -w production -p support-agent

For a tool call, verify:

  • tool_id exactly matches a registered tool;
  • the EMU includes required tool.version;
  • the intended tool registration and runtime executor are connected to this project; do not infer that from a validator result that may use broader registry fallbacks;
  • the runtime dispatcher recognizes the action type and version;
  • credentials and network access exist in the executor environment;
  • errors are surfaced rather than swallowed;
  • the activation is acknowledged only after the intended outcome;
  • success emits a material event.

For routes and decision prompts, inspect the destination registry and timeout behavior. For context directives, confirm they are actually added to the next prompt rather than merely logged.

Registration returns 422#

Check the enforced schema rules:

  • every tool_call.tool includes "version": "1.0.0" or the actual registered version;
  • lifecycle state is lowercase;
  • cooldown is { "seconds": N, "gate": "ack" } or uses a supported gate, not an ISO string;
  • idempotency is a structured object;
  • state keys are lowercase and namespaced;
  • the authorization scheme for direct HTTP is AMI-Key, not Bearer.

Run memrail emu-plan ./emus/ --strict before applying. Strict mode validates candidate definitions against the server and rejects errors, unavailable validation, or incomplete reports before writes. See Production workflow for release checks.

Trigger syntax is rejected#

Common causes:

Memrail DSL
// Invalid: lowercase logical operator and double quotes
state.user.tier == "premium" and tag.intent == "upgrade"

// Valid
state.user.tier == 'premium' AND tag.intent == 'upgrade'

Also check the mandatory IN 'duration' clause, three-token subject.verb.object event names, numeric comparisons against numeric state, and template literals in event WHERE filters. A duration beyond configured retention is syntactically valid but may be flagged as a temporal-feasibility warning.

See the Trigger DSL reference.

A model tag creates inconsistent behavior#

Inspect the boundary before Memrail:

  • Is the classifier output constrained to a fixed enum?
  • Is casing normalized before building the tag?
  • Is unknown a supported value?
  • Is the tag always produced or conditionally omitted?
  • Did a model or prompt version change the distribution?

Do not solve classifier drift by adding free-form variants to deterministic policy. Restore the contract at the inference boundary.

Wrong project or workspace#

Print non-secret scope values at startup and compare them with CLI arguments. Default team behavior can hide an unexpected scope if one service sets AMI_TEAM and another does not.

bash
env | rg '^AMI_(ORG|TEAM|WORKSPACE|PROJECT|BASE_URL)='
memrail list-emus -w production -p support-agent

Never print AMI_API_KEY.

Keep the EMU, intended tool registration, and executor coherently scoped to the same project as an integration contract; current registry fallback behavior is not a hard guarantee of that coherence. Cross-project event queries are separate and should use an intentional documented dependency.

Diagnostic codes#

Download the error, warning, and suppression catalog. It separates EMU validation findings, decision trace suppressions, and HTTP error codes, so agents can branch on identifiers instead of parsing English messages. Not every HTTP error has a code; always retain its status and response details.

For invalid EMU structure, validate against the EMU JSON Schema, then use strict preflight for DSL and registry checks. The catalog includes suggested responses; a warning is evidence to investigate, not automatic permission to change policy.

Evaluation changes live behavior#

Check custom dispatch against the evaluation and execution contract, and cohort configuration against the lifecycle contract. Use production rollout checks to test those boundaries.

Acknowledgment returns 404#

Check the activation ID, organization/workspace/project scope, and receipt eligibility under the outcome and acknowledgment contract. Retry transient delivery failures with bounded backoff.

Lock file is out of sync#

Pull the current remote state, inspect the change, then resolve source differences:

bash
memrail emu-pull ./emus/ -w production -p support-agent
memrail emu-plan ./emus/ -w production -p support-agent --strict

Do not manually forge lock hashes. If remote policy changed outside the reviewed workflow, preserve that evidence and reconcile it explicitly.

Unsafe fallback detected#

An agent control call that fails and then executes the original model proposal is a bypass. Define failure by decision-point consequence:

Decision point Reasonable fallback
prompt steering base prompt without optional directive
read-only retrieval bounded public corpus or no result
external tool with side effects deny or require human review
irreversible workflow transition hold current state
notification queue for retry with idempotency

Test timeout, authentication failure, malformed response, and empty selection. Fallback is part of the control policy even when implemented in application code.

Reset a development workspace#

Workspace purge is destructive and requires an organization-level key:

bash
memrail purge-workspace development --targets emus,traces,events --yes

Confirm the exact workspace and target list before running it. Do not use a production example as a copy-paste default.

Minimal diagnostic report#

When escalating an issue, include:

  • SDK and CLI versions;
  • non-secret org/workspace/project and decision-point names;
  • EMU key and lifecycle state;
  • trigger and declared action type;
  • redacted context keys with value types;
  • validation warning codes;
  • invocation or trace ID;
  • whether selection, dispatch, acknowledgement, and outcome event occurred;
  • expected and actual behavior.

Do not include API keys, personal data, or raw sensitive prompts.