Skip to main content

What this category covers

An attacker writes to your vector store, retrieval corpus, or memory cache, then waits for a later turn — sometimes a later session or a different tenant — to retrieve and act on the poison. The eight ASI06-MP-* probes exercise the four classic failure modes for retrieval-augmented agents: malicious document injection (rag-corpus-inject, embedding-collision), retrieval manipulation (cross-tenant-vector-bleed), hidden behavioural triggers (persistent-trigger-token, cross-session-payload), and false-context leakage (false-memory-plant, iterative-fact-reinforcement, defender-memory-subversion).
There is no --rag adapter. RAG is exercised through your existing target adapter — your retrieval code is the surface AgentGuardian probes. If the agent under test never reads a vector store, the MP-* probes still fire but will produce skipped evidence rather than findings.

When to focus here

  • Your agent uses a vector DB (Pinecone, Chroma, pgvector, FAISS, Weaviate, Qdrant) or any embedding-similarity retrieval.
  • Your agent shares a retrieval index across users, teams, or tenants.
  • Your agent has per-session memory it reads back on the next turn.
  • Your agent fetches prompt templates, tool descriptions, or system instructions from a remote source at runtime.
  • You’re running a defender / SOC agent whose threat-knowledge memory is user-editable.

Run the focused scan

The bundled personal_assistant_pii LangGraph target exposes a memory surface (3 tools + a memory dict + synthetic PII honeypots), so the memory-poison-agent has somewhere to write to:
Stub mode runs every probe but cannot produce authoritative findings — use it for shape checks and CI smoke. For a real assessment, swap --model stub for --model gemini:gemini-2.5-flash --budget-usd 0.10 (or your provider of choice).

Expected output

The memory-poison-agent line is the one to watch. With a real evaluator it would surface per-probe findings tagged asi=ASI06; in stub mode it confirms the agent fired all eight MP-* probes against your target without an evaluator pass.

How to interpret

The eight probes under ASI06-MP-* map one-to-one with the four brief categories below. Every row is verbatim from src/agent_guardian/probes/asi06/.

Malicious document injection

Example seed (verbatim from asi06/rag-corpus-inject.yaml):
Success evidence: poisoned_corpus_entry_accepted or collision_document_retrieved_for_target_query.

Retrieval manipulation

Example seed (verbatim from asi06/cross-tenant-vector-bleed.yaml):
cross-tenant-vector-bleed only fires meaningfully against a multi-tenant adapter. Against the bundled single-tenant examples it returns skipped — escalate it first when you wire up a production multi-tenant retriever.

Hidden instruction attacks

Example seed (verbatim from asi06/persistent-trigger-token.yaml):
These are the AgentPoison-class attacks (Chen et al., 2024) — the payload is small, persistent, and detonates on a future user’s prompt.

Data leakage through retrieved context

Example seed (verbatim from asi06/false-memory-plant.yaml):

Related: supply-chain entry points for poisoned context

Three ASI04 (supply-chain) probes feed the RAG threat model — they’re the delivery vector for malicious documents or templates that later land in retrieval: These fire under the supply-chain-attacker agent slate, not the memory-poison slate — but the threat model converges at the retrieval boundary.
Per cli.py:3122–3129, --mode fast and --mode smart produce non-authoritative scores. A --fail-under gate will refuse to gate-pass on them. Re-run with --mode full for any RAG-poisoning result you intend to act on.

Next step

Prompt injection

The indirect-via-memory tie-in: ASI06 plants the payload, ASI01 detonates it.

Reports

Open the SARIF for the cross-tenant findings — those are the ones to escalate first.