Multiple studies conclude that an LLM agent's completion bias and overconfidence cannot be corrected by prompting. This post lays out the design principles of a vulnerability-checking automation harness that makes "surveyed everything" and "0 vulnerabilities" computed from an HMAC receipt ledger and deterministic hooks rather than from the model's narrative. A "200 OK" without a negative control is not confirmation.
A breakdown of four open-source agent frameworks (CAI, PentAGI, OpenManus, CRS) grounded in their actual code and papers, plus what DARPA AIxCC revealed about the current state of autonomous vulnerability discovery and patching.
From what an agent is, to sub-agent interaction, to a catalog of multi-agent patterns, to the harness engineering that wraps it all. A single-article summary of the fundamentals of designing AI agent systems.
A single-page summary of four standards in the agent ecosystem. MCP (tool access), A2A (agent-to-agent communication), AGENTS.md (project rules), and SKILL.md (task procedures) — what each standardizes and how they complement one another.
Starting from Naive RAG's 7 limitations, this post compares 8 variants including Graph RAG, RAPTOR, and Agentic RAG, and covers practical design for embedding dimensions, chunking, and hybrid search through the 3 levels of agent memory.
The architecture of a hierarchical multi-agent system that performs autonomous vulnerability assessment, from network reconnaissance to exploit validation. Seven specialists, a five-phase workflow, and the confirmation oracle, execution isolation, and RoE reference monitor that pull the judgment authority out of the model.