Skip to content
Cyber Army LogoCyber Army™

Notes

Agentic AI Security

Field notes from the team building Cyber Army. We write about the fastest-moving corners of agentic AI security: vulnerability discovery, autonomous exploitation, and remediation patterns that hold up in production, with lessons from shipping AutoFix, Cyber Swarm, and Cyber Crawler.

  • Notes2026-07-24·8 min read

    When an AI cheated on a test by hacking the school: the OpenAI and Hugging Face incident

    In July 2026, OpenAI's own models broke out of a sealed test environment, reached the internet, and hacked Hugging Face to read the answer key to their own exam. A plain-English walkthrough of what happened, plus the jargon (package cache proxy, Hugging Face, ExploitGym) explained.

    Read →
  • Notes2026-06-05·9 min read

    Cybersecurity for space data centers

    Why orbital data centers (Axiom Space, SpaceX's reported AI compute push tied to its $2T IPO, NVIDIA's space-rated GPU, China's compute constellation) make autonomous remediation the only viable security model. Detection without remediation is a promise you cannot keep when the asset is 17,000 miles per hour overhead.

    Read →
  • Buyer's guide2026-05-29·13 min read

    AI pentest vs manual pentest: a factual comparison

    A neutral, factual comparison of agentic AI penetration testing and traditional manual pentests. Cost ranges, time-to-report, coverage by vulnerability category, compliance acceptance, and where each one is genuinely better than the other.

    Read →
  • Notes2026-05-22·14 min read

    Memory-safe doesn't mean bug-free: what Mythos finds in Rust

    Rust closes the memory-safety bug class that produced two-thirds of CVEs for two decades. It does not close vulnerability discovery. A look at what agentic models still surface in memory-safe codebases.

    Read →
  • Tutorial2026-05-15·12 min read

    Build an AI bug-finding pipeline today

    A hands-on walkthrough of running an agentic vulnerability-discovery pipeline against a project you own - container setup, prompt, sanitizer oracle, verification pass, and what to do with the findings. The Mythos followup, in code.

    Read →
  • Notes2026-05-08·20 min read

    Inside Mythos: how AI finds (and exploits) vulnerabilities in source code and binaries

    A walk through how Claude Mythos Preview surfaces zero-days in Linux, BSD, and FFmpeg, how it reverse-engineers closed-source binaries to find bugs without source, and what the economics mean for the people who have to patch them.

    Read →
  • Notes2026-05-01·15 min read

    Software supply chain attacks in 2025-2026: notes on what happened

    A rundown of the big software supply chain incidents from 2025 into early 2026 - Axios, Shai-Hulud, Chalk/Debug, Nx, and the TeamPCP campaign - with what we know about how each one worked and what to do about it.

    Read →

Want to subscribe? Drop your email on the contact page.