logoalt Hacker News

500B tokens later: Letting AI agents decompile a first-person shooter

136 points • by davikr • today at 2:02 AM • 102 comments • view on HN

Comments

brandonpelfrey • today at 2:32 AM

Unless you explicitly need byte-matching decompilation, there are significantly faster ways to produce a decompilation/C which is functionally equivalent. I need to post about this. What's been working for me is that for every function, Agent A is tasked with writing some code which is semantically equivalent to the original assembly, but not necessarily exactly the same. Agent A also writes tests. Agent A submits the implementation of the function and tests to the harness for it to judge. The harness runs both the original function and the submitted function in a virtual machine/simulator/emulator (the tests define function inputs and starting state). The harness will only accept the implementation if 1) the read/write sequence to RAM is identical to the original function's, and 2) there must be complete line and branch coverage of the original function being decompiled.

I've found this to be robust for decompiling games, while giving the agents enough freedom to write code that is readable and not waste a ton of time making sure e.g. instruction ordering, register assignments, etc. are all exactly the same. For me, having byte-matching decompilation is only one way to produce a decompilation I know is faithful to the original. This "high-level decompilation" process I just described is something agents can do much more quickly.

➕ show 3 replies
ChoosesBarbecue • today at 11:10 PM

It’s probably worth noting that the author is decompiling a game that he’s had ~10 years of experience modding. And has been involved in finding vulnerabilities in [0].

[0]: https://github.com/momo5502/cod-exploits

➕ show 1 reply
aetherspawn • today at 3:11 AM

The reason this cost so much is because the AI has the ridiculous goal of getting identical assembly output.

The agents would have had to mess around with compiler versions, optimisation options, and the phase of the moon as well.

If you just went for functional equivalence, it would probably cost 10x or 100x less tokens.

Another false economy was using Sonnet instead of a more intelligent model like Sol 6.1 (1), which would have cost more per token, but is 100x or so better at reverse engineering and coding and therefore can chew through the source code much quicker and make fewer mistakes, meaning less work needing to be scrapped.

In my testing doing a similar task, I ran multiple sonnet for weeks and burnt through ~$1000 in tokens to get 20% completion and output that was pretty bad. After switching to Sol 6.1, it finished the whole task in around 2 days, cost around $50, and it did it with zero supervision and a single /goal.

(1): struggle to use Opus for reverse engineering, too many safeguards. OAI has virtually none, and uses way less tokens so is more economical.

➕ show 2 replies
nvme0n1p1 • today at 3:36 AM

> The avid reader of my blog might have noticed that I had previously written two posts that have since been removed. Everyone else might now be wondering which game I am talking about. To both of you I can only say that corporate America was here to ruin our fun.

Call of Duty: Modern Warfare 2 (2009)

https://web.archive.org/web/20260925153118/https://momo5502....

https://web.archive.org/web/20260925153131/https://momo5502....

Come at me, corporate America.

➕ show 2 replies
jakubmazanec • today at 11:17 PM

> Architectural decisions were not questioned, as long as they aligned with the goal.

Yes, that's the main problem I'm having with agentic workflows. Sometimes you don't care about architecture of something if it just works, but more complex stuff can't work well without good architecture. And currently it seems that there is still no substitute for good human taste.

stevefan1999 • today at 1:09 PM

Reverse engineering dynamic dispatches like vtable and fat pointers/slices and traits would be hell. Different padding and compiler settings also contribute a lot.

One of the few problem is that the decompiler is not always reliable. That forces you to go read the assembly, and it is not apparent to decipher the right kind of feng shui, not even from human before the LLM era. I used to play CTFs and my conclusion is exactly that.

This is extremely apparent when there are self-modifying code (e.g. JIT) is involved. You need a stepping debugger to read the right control flow, because the code will diverge based on the instruction pointer and regions you jumped into. At this point static analysis like IDA and Ghidra stopped working.

edg5000 • today at 3:41 AM

I sense the approach overcomplicates things. I wonder how long this would have taken in a single session. Maybe this is actually a textbook example of something where subagents make sense, but when I first started LLMs I was often overcomplicating the workflow with all kinds of orchestration. Now I just use one agent, it better allows controlling the output even if the agent works slightly longer. Most time is spent by me writing prompts and reviewing work anyway (for me at least).

➕ show 1 reply
testerius • today at 4:42 PM

Is it the same Call of Duty: Modern Warfare 2 decompilation project? If yes, then why author says he/she does no release source code? If I remember correctly, it has to be open sourced Modern Warfare 2 modernized engine and so on? Or do I miss something?

➕ show 1 reply
WheelsAtLarge • today at 3:09 AM

Interesting, if all software can be decompiled and copied what is the future of software. Will all software be SaaS? A time where the majority of PCs will be terminals? Game consoles are almost there. It's only a small jump for all software to go that way.

Edit: Here's a possibility.

The future of software is agent only software. We ask for a result, agent asks questions from us, agent uses the specialized software, user gets result. We subscribe to an AI assistant and specialized agents. We are almost there,at least the start. The future of PC's as we know them are numbered. OSs,CLI,compilers and whatever will melt into AI assistants. Say goodbye to writing software for people.

➕ show 4 replies
bob1029 • today at 8:16 AM

Reverse engineering existing game binaries seems kind of silly when a lot of the most important AAA knowledge is already available to the general public.

https://github.com/ValveSoftware/halflife/blob/master/pm_sha...

The LLMs have presumably already consumed information like this as part of their training sets. Converting between Hammer and Unity scale is a fairly trivial linear operation. You can dump the BSPs to recover geometry and rapidly accelerate map development using timings that are already known to work.

Attacking the raw binary directly is certainly impressive, but it's totally unnecessary.

petetnt • today at 8:40 AM

Great article that really demonstrates why the copious amounts of AI generated PC ports do not fit under ”preservation” guise, which to me has been the point of doing these ports in the first place and what many of the enthusiasts before the LLM wave were aiming for.

> There are no noticeable bugs and all features of the original game are present.

>

> The remaining functions have been reworked repeatedly. While they still don’t match byte for byte, we believe their semantics are correct.

As there are no way to confirm this outside of ”believinh”, what you are left is with a end product that might contain thousands of micro changes that essentially make the game something else than it was intended to be.

➕ show 1 reply
thway15269037 • today at 2:55 AM

I struggle to understand what legal leverage they used to threaten him to remove every detail about the game. Can someone post the game name and company name?

So, if you reverse-engineer game X and post reverse-engineered code, what exactly do you infringe, how and in which jurisdiction? What changes if it is done via LLM?

(I understand that LLM decompilation is absolutely out of hand right now and something surely will come to trample the fun. But what and when? I suppose american LLMs will have their system prompt updated to forbid any reversing help and report suspicious activity straight to legal hotline)

➕ show 3 replies
Neywiny • today at 12:39 PM

I do wonder if LTO would obfuscate this a bit more. Presumably if you can't chunk it into small compilation units, you can't find a way to make C from the machine code.

mawadev • today at 6:41 AM

It is wild how AI ends up in the trenches of plagiarism, going from art to text, from software to games...

otagekki • today at 8:32 PM

what if someone does this to the Windows OS?

Fizz43 • today at 9:57 AM

the 80% is a meaningless AI generated % by the way.

georgemcbay • today at 2:56 AM

In case anyone is curious about the obvious question of which game they are talking about... based on months-old reddit posts (which seem like links to prior progress reports of the same project) the game in question appears to be Call of Duty: Modern Warfare 2 (the original 2009 version).

https://www.reddit.com/r/ReverseEngineering/comments/1vxig19...

xnx • today at 2:58 AM

A previous version of the page said the game was Call of Duty: Modern Warfare 2 (2009).

vivzkestrel • today at 3:22 AM

- i want to very very badly see a post of this using LM studio and one of the open source models

- please someone do it

tasubotadas • today at 8:40 AM

Interesting read and a thread here. I am working on decompiling two obscure abandoware games from my childhood and it's been a difficult and quirky process...

Even after building whole infra for that and not pursuing identical byte output it's been pretty token hungry process

Uptrenda • today at 9:03 AM

I'm more interested in how the OP managed to convince claude to do what was probably against Activison's TOS. I am assuming they didn't hand specifics of the game but just a blank binary. Because if you had of given specifics claude would have checked the TOS of the IP holder, found that the company doesn't want people to decompile their software, and refused to work on it. Or am I missing something here?

esafak • today at 3:22 AM

Can anyone think of any lessons to draw from this for normal development, where we don't have oracles to serve as guardrails? I write specs but the agents still find ways to insert bugs between the lines. Oh well, job security.

➕ show 1 reply
carplaybits • today at 2:11 AM

[flagged]

WillAdams • today at 2:29 AM

To save folks looking this up:

>At current 2026 API rates, 500 billion AI tokens would cost roughly $100,000–$750,000 depending on the model, with most flagship models in the $150–$400 per million input tokens range

➕ show 4 replies
FounderDenken • today at 3:11 AM

[flagged]

danielsorok • today at 8:48 AM

[flagged]