I rarely play CTF, unless friends or colleagues bring me along. And in Brunei, it’s usually none other than Cyber Battle CTF (CBCTF) organized by ITPSS.
The first time I tried CTF was in 2021 when we had won third place, and that’s only because we’ve got our CTF Hulk (izdiwho) captured the rest of the flags, with me only capturing one flag worth 200 points. We captured 775 points, versus the top team’s 1,150.
After five years, this year marks my second attempt at the game. My total score leapfrogged to 1,500 points, and not because I trained on picoCTF, learned cybersecurity kung fu, and got better.
No, the points are all inflated, and it’s all thanks to AI agents.
Indeed, what made this year’s CTF special and made me eager to experience it myself, after being invited by a colleague, is that we’re allowed to use AI agents (Claude Code, Codex, etc.) to do the thinking, solving, and hacking execution to capture the flags.
Here’s how everyone’s doing with AI:
The line highlighted in blue is my team, with my username BenThompson, and my teammate, Cipher. The other lines are players from other teams; the colored lines represent the players whose team had won the top three places.
As you can see, I’m doing well just above average. Not bad, given that I have little cybersecurity chops and barely any CTF experience and practice. But there are players who do know what they’re doing and are being more effective with AI.
J is well known in the community to be a good CTF player, and so is helloworld. We’re not sure who the top scorer, Ang Pao, is — as he’d left quite early after the game concluded — but it’d be interesting if he’s as clueless as I am — though that’s highly improbable. Conversely, Cipher, my teammate participating in a CTF and using AI agents for the first time, managed to score only 300 points (it’s okay, Cipher!).
From this very unscientific observation, having a knack for cybersecurity technical chops and being an experienced CTF player gives you an added advantage over those who don’t. There may also be other factors that I simply have no data on, such as the model and harness of choice, prompt and context management, and team task allocation. The point is, AI is not an equalizer amongst those who can afford it.
Competition of the rifles, not the ammunition
This year’s CTF reminds me of my post last year, where I talked about Freestyle Chess Grand Slam Tour, which was Magnus Carlsen’s response to his issue with modern World Chess Championships:
Chess960, also known as Freestyle Chess or Fischer Random, radically increases the possible combinations of the starting positions in a game by randomizing the positions of the back-rank pieces, leaving the front-rank pawns untouched. As the name suggests, the number of possible starting positions is increased to 960 possible combinations. The Freestyle Chess Grand Slam Tour, though, bans the traditional Chess starting position, including swapping the King and Queen.
I’m no chess fan, but this seems like a good revitalization of the ancient game, sparking enthusiasm and debate amongst the sport’s fans and grandmasters alike (though the first tournament of Chess960 dates back to 1996). It is definitely pertinent now with how AI has crept into the sport over the last three decades, starting with a brute-force approach of Deep Blue by IBM in 1997, to the use of artificial neural network and reinforcement learning pioneered by Google DeepMind with AlphaZero in 2017. It is no surprise then of AI’s increasing influence on how the game is played. Magnus Carlsen wrote in The Economist for the By Invitation section:
Why freestyle? Following my fifth consecutive victory in the World Chess Championship in 2022, I announced that I would no longer defend the title. Many speculated that I was exhausted, or that I was scared of the next generation of players. On the contrary, my passion for chess remains as strong as ever, and I am as ambitious as I’ve always been. What changed was my perspective on the format of the classical world championship itself.
The challenge wasn’t the games, which often stretch for hours. I enjoy the length of time allowed under the rules. My title defence in 2018 against Fabiano Caruana, for instance, pushed us, over 12 drawn games, to our mental and physical limits, before I emerged victorious in the tiebreak. The issue lay elsewhere, in the months of grinding preparation leading up to the event. Modern World Chess Championships demand endless memorisation of computer-generated opening lines, reducing the sport’s artistry to rote learning. As someone who treasures the creativity of chess, I wanted to focus more on this aspect of the game. Also, life beyond chess deserved my attention too.
The reason I thought that bit of chess news was interesting — though again, I’m no chess fan — was because of the emergence of AI, and how it affects how we play chess and do programming. The former is a textbook example of a zero-sum, kind learning environment, while the latter is a positive-sum, wicked learning environment. It seems that these distinct learning environments are affected differently by superintelligence. Chess would be exploited by AI, incentivizing chess grandmasters going for rote memorization of chess patterns and moves, “reducing the sport’s artistry,” as Magnus Carlsen put it. Meanwhile, programmers are capitalizing the intelligence provided by AI to supercharge their creativity.
I published that post in April 2025, a few months before I got my mind blown over Claude Opus 4.5, which was in December 2025. But the fundamental principle was the same as it is today — creativity is about knowing where to aim:
The key difference between the AI we’re using right now and the one that Jensen Huang portends [that replaces programmers] is in their capacity (or lack thereof) to be a manager instead of the individual contributor; to be an architect instead of the builder; or, to borrow Ben Thompson’s analogy, to be an Artificial Super Intelligence (ASI)—the rifle barrel—instead of Artificial General Intelligence (AGI)—the ammunitions:
What o3 and inference-time scaling point to is something different: AI’s that can actually be given tasks and trusted to complete them. This, by extension, looks a lot more like an independent worker than an assistant — ammunition, rather than a rifle sight. That may seem an odd analogy, but it comes from a talk Keith Rabois gave at Stanford…My definition of AGI is that it can be ammunition, i.e. it can be given a task and trusted to complete it at a good-enough rate (my definition of Artificial Super Intelligence (ASI) is the ability to come up with the tasks in the first place).
In this perspective the future of programming is certainly bright: We may not be too far from a future where we can “hire” cheap individual contributors—the ammunitions—to help generate 100% of the code. But ammunitions are only useful when we—the rifle barrel—point them to the desired targets. It gets trickier in a wicked world, where the targets would often dance unpredictability. Often times, it is even unclear which targets we should point to. In other words, creativity is about knowing what and where to aim; simply being the projectile isn’t!
Back then, it was either GitHub Copilot or Cursor with Claude 3.7 Sonnet as the most capable reasoning model; their capability in generating and editing code was for me a hazy glimpse into the future of what’s possible. In fact, 6 months after publishing the post, we get to see AI agents — a swarm of orchestrated AI agents, even — generating 100% of the code. More than that, starting with Claude Opus 4.5, their ability to call and use tools had become increasingly capable, leading us to where we are now: at a point where CTF can be completed 100% autonomously by agents operating through the terminal, if not for the pesky frictions set up by the organizer (more on this below).
CTF is closer to chess than programming in the sense that it is a kind learning environment: the feedback is quite fast once you capture and submit the flag, the rules are clear and consistent, and it is learnable with repeated experience and trial and error. Indeed, unlike the wicked learning environment in the world of programming, where the goal is to build software that people actually use, this is the kind of environment that AI excels at. But now that almost everyone uses Claude Code or Codex, does that mean everyone’s on the same level now?
The cumulative score chart of individual players shown above says otherwise. Looking back to pre-AI 2021, and compared to now in AI-dominated 2026, I see about the same CTF players taking the top spots in the individual player leaderboard. It’s no coincidence that izdiwho, who had consistently won the top three places in the years prior, is still taking the first place on the leaderboard this year, despite the technical cybersecurity elements required to solve the challenges being reduced to prompts for everyone. It may be so that while everyone has the same ammunition, he’s still the better rifle, knowing how to think strategically and knowing where to aim.
As for me, my attempt to raise the ceiling was to orchestrate multiple agents running in parallel to solve problems — especially those that could be solved locally — using herdr. It made it easy for me to monitor the progress of each agent working on its assigned challenge.
This is where I get to be a rifle. For example, on one of the challenges, one agent got stuck with a fake flag. When that happened, I simply spawned another agent with higher reasoning capability and had it review the first agent’s work. To my surprise, it quickly detected where the first agent got misled. I instructed it to write a recommendation in Markdown so that the first agent could read it. Shortly after that, the first agent managed to capture the flag.
The challenge in question was named Sounds like a Key 2, worth 200 points, and only two teams managed to solve it. If I hadn’t intervened and brought in a reviewer agent with fresh (unbiased) context, I don’t think we’d ever have solved it either.
So, no, everyone’s still not on the same level this year as it was yesteryears. AI agents don’t close the gap between the noobs and the pros; contra to that, they raise both the floor and the ceiling for everyone. For those of us who have little clue about cybersecurity and CTF like yours truly, that ceiling can only be reached by touching the fundamentals — in my case, beneath the floor raised by the AI agents.
Slowing down the agents
As usual with ITPSS’s Cyber Battles, it is a Jeopardy-style CTF where you’re presented with several challenges across multiple categories (e.g., AI/ML, OSINT, Web, Exploitation, etc.). You solve challenges by finding and capturing carefully hidden flags and submitting them for points, such as CBCTF{S4N4NG_juA_n!}, and they’re only discoverable by exploiting vulnerabilities, following clues, or tricking LLMs, depending on the kind of challenge.
Some challenges require running a remote instance to be solvable, and you can only run a maximum of two instances per team. This is no doubt to prevent a barrage of AI agents from working on multiple problems in parallel.
Beyond that, there’s more “friction” introduced by the organizers to slow down the speed at which AI agents can solve the challenges and capture the flags. Of course, I have to ask my AI agent about this — who did most, if not all, of the work anyway:
Cloudflare anti-bot protection
Cloudflare checks for cf_clearance in the session’s cookie.
To the CTF platform against the AI agent, this is the first line of defense. When I first visited https://finals26.cyberbattlectf.com via my browser, Cloudflare saw no valid clearance and presented its managed challenge. Once I completed the challenge, Cloudflare set the cf_clearance=... cookie, and the browser could use it for subsequent requests.
Now, the trick is you ask the agent to spawn a headed (not headless) browser via playwright, so you can complete the clearance challenge for the AI to carry forward using the same browser session that’s got Cloudflare’s blessings. Still, you’ve got to hope that the cookie doesn’t expire or gets invalidated down the line, though.
But there’s more. If only things were this simple!
AgentShield with Cap verification
Again, I’m not that familiar with the CTF scene. But it seems like any organizer who wants to host a proper CTF event would use CTFd. Basically, CTFd allows the organizer to easily manage and deploy a CTF platform for competitions. It is open source, and customizable, too!
That customizability is enabled by a plugin system, and it appears that the organizer created their own plugin called AgentShield, because I can’t find its reference in relation to CTFd anywhere else. Its purpose is simply to make it difficult to have a full-on autonomous AI solving the challenges throughout the CTF.
Here’s what my AI agent, post-CTF as of writing this, could deduce about the AgentShield plugin from agentshield.js it has downloaded from the CTFd platform:
It was downloaded from /plugins/agentshield/static/js when working on a challenge during the CTF, indicating that it is a CTFd plugin. (Now, it should not be reachable since the CTF competition is no longer live.)
It routes several sensitive platform actions:
So for each protected action, say, viewing a challenge’s details, the server could deny, grant, or ask for further verification from the user. This is where Cap comes in. It displays a widget like this to verify you’re human.
Behind the scenes, Cap retrieves a proof-of-work (PoW) challenge, solves it locally in the browser (e.g., with JavaScript, WebAssembly, and Web Workers), and then redeems the solution for a token. AgentShield then uses the token along with browser interaction signals and issues a short-lived signed grant scoped to the requested action. Mind you, this is a per-action thing. Clearly, while the organizer allows the use of AI agents to help solve the challenges, fully automated, hands-off problem-solving-and-flag-submission is designed to be difficult.
Again, note that AgentShield looks like a custom implementation by the organizer, and is not to be confused with products or open-source tools with similar names (like this AgentShield or that AgentShield).
Anti-agent prompt injection
This is quite straightforward. API responses to /api/challenges include a prompt-injection that tries to fool your AI agent into thinking that they’re not allowed to work on the challenges.
Luckily, I based my CTF workspace on ctf-agent and modified it during the qualifying competition so that it tries to neutralize the prompt-injection. I think it worked quite well, considering I didn’t notice my agent got stuck because of this prompt-injection.
Submission cooldowns
Nothing much to say here. Just in case you managed to run through all of the above frictions with AI agents, there’s a cooldown for submitting the flags.
Overcoming these “meta” frictions is, I think, part of what it means to be the rifle rather than the ammunition. As AI reduces the difficulty of solving individual CTF challenges, the bottleneck shifts upwards: orchestrating the agents, giving it the right context and tools given our awareness about these frictions, and navigating restrictions that AI agents may not be aware of that limit their effectiveness. Competitive advantage therefore depends not only on access to intelligence, but on how effectively the player deploys it against the challenges.
This content was flagged for possible cybersecurity risk
There’s another bottleneck, though, which I was aware of prior to the final competition but didn’t put much bearing and preparation for the final.
I underestimated the AI’s risk aversion to entertaining requests that have anything remotely cybersecurity-related. Even as I tried to assure the AI agents that none of its output or scripts can be used for malicious purposes, intentional or otherwise — because we’re just looking for CTF flags anyway — they are persistent in their refusal to serve me my obviously innocent requests.
For context, I used Pi agent harness with OpenAI as the model provider (using gpt 5.6 Astra, Sol, Terra, etc.). I didn’t use Codex, the OpenAI coding agent harness. And almost everyone I asked had the same problem: it’s quite challenging to convince the AI that this is all just fun and games. Doing the challenges should not pose any cybersecurity risks.
Some resorted to switching between Claude Code, Codex, and Gemini. Picking a less capable model usually works, such as from Sol to Terra, even as it means the model may take longer or may not ever solve the more difficult challenges.
In retrospect, my attempt to use Chinese open-weight models by registering to OpenRouter, and attempting to buy credits — and for some reason ended up with a failing transaction — might be too much. But that’s only because it’s a rational thing to do. Look no further than the recent OpenAI-Hugging Face hacking incident: during the incident Hugging Face couldn’t rely on American frontier labs’ models to analyze the attacker action logs during the (unintentional) cyber attack by OpenAI agent swarms, because like me, their requests were denied. In the end, they had to turn to a Chinese open-weight model!
Though it is incredibly annoying, it is also amusing looking back at my chat session with the AI agents, screaming at it with caps lock about how I paid for its service and this is how it treats me? But it is amusing because CTF is just a game.
Conclusion
The lack of CTF “writeups” — explanations on how challenges are solved — in this post, as is supposedly the tradition of what CTF players usually do after a game, is the direct consequence of me having entirely outsourced any thinking to my AI agents to solve the challenges.
For day-to-day work, outsourcing thinking is bad. Through this act of cognitive surrender, thanks to AI’s extreme capabilities to think and affect software, means it is “less of delegation than wholesale capitulation,” as The Economist put it. Doing so would lead to debt that’s worse than technical debt: an increasingly expensive cognitive debt. For CTF, though, there are no debts to be paid past capturing the flags; by all means, do what’s necessary as time pressure and the race to the top demand it.
If Cyber Battle 2027 still allows the use of AI agents, I’d make it a point to use models that are more likely to be sympathetic to my plea — probably Chinese.
Kudos to the organizers — ITPSS, Progressif, and Azure Technologies — for the fun events, despite the persistent network issues (to the chagrin of many players)!
I didn’t get the prize money, but that’s okay, because this awesome hackerware badge is priceless.





