Article in progress
This article covers the direct attack surface: prompt injection, jailbreaks, goal hijacking, and system prompt extraction via the input channel. Defense means trust boundaries, I/O validation, and session isolation inside the LLM application itself.
Topics will include: prompt injection taxonomy, jailbreak mechanics, multi-turn goal hijacking, system prompt extraction, guardrail bypass patterns, and application-layer defense primitives.
In the meantime, read From Conversation to Action for the full attack surface overview, or Threat Modeling: Necessary and Sufficient Conditions for the first-principles framework.