logoalt Hacker News

lonlundgrentoday at 7:08 PM1 replyview on HN

Thank you for your kind words, Aurornis.

Yes, this is my production corpus, across 65 usage days, two subscription accounts, three machines, 25 project groups, and 213 sessions, across 43,261 invocations and 7,583 turns. Use your own data if you want to prove or refute what was seen in my corpus.

The "ups and downs in the chart" were not plotted with sub-daily resolution. Specifically, the two-month temporal chart uses a 3.5-day Gaussian bandwidth which is meant to reduce short-term noise while retaining broader changes.

Additionally, a separate episodic analysis identified multi-day changes in delivered thinking. And those episodes were predictive of held-out work. The more projects pulled into an ensemble, the more predictive they were of delivered thinking tokens for held-out projects during the episode.

The point of the benchmarks is to establish a baseline for what thinking-token counts one should expect from specific effort levels using published numbers, not what every response should be delivered. If you looked closer, you would see that P90 invocations were still delivered 13x thinking tokens below that level.

So you don't have to provide a generous interpretation of my workload if you don't want to. Remove all of the zero-thinking token responses, redistribute those samples across the distribution, and then tell me if it magically shifts right and starts delivering anything close to published numbers. Only -46- out of 36,374 July and August invocations even broke 16k thinking tokens - only 0.13% of the total. You are welcome to present what percentile you think is a fair comparison to make here.

If you would like to denigrate a month of my time as vibe-slop, that is your prerogative. You can even be dismissive of my workload, if you want, even if my background should tell you otherwise. But if you want to knock a month of someone's time, do it with your own data to at least help move the conversation forward.


Replies

Aurornistoday at 7:16 PM

> Use your own data if you want to prove or refute what was seen in my corpus.

I don't think you understand. What you posted is highly dependent on your corpus. I can't "refute" anything because it's not available and it's the major variable in the experiment.

> The point of the benchmarks to establish a baseline for what thinking-token counts one should expect from specific effort levels using published, not what every response should be delivered. If you looked closer, you would see that P90 invocations were still delivered 13x thinking tokens below that level.

I think you're missing something from that first sentence, but I assume you're talking about the comparison to ARC-AGI-2 published thinking tokens?

It should be blindingly obvious that you do not want your thinking token counts to be as high as a benchmark that was designed to push LLMs to their limit.

> Only -46- out of 36,374 July and August invocations even broke 16k thinking tokens - only 0.13% of the total. You are welcome to present what percentile you think is a fair comparison to make here.

What point are you even trying to make?

Again, you don't want invocations to be burning 16K thinking tokens except for rare problems that 1) must be solved in one step and 2) are designed to be entirely self-contained thinking in that step.

You're trying to compare development work to a benchmark that encapsulates complex thinking into a single step.

Coding work is iterative and works in incremental steps: It runs commands, reads more files, checks the web. Thinking tokens should be low for your turns.

ARC-AGI problems have an input and an output. They look like this: https://arcprize.org/tasks/b5ca7ac4 They have more thinking tokens because that's the entire state. They get one output and it's constrained.

> But if you want to knock a month of someone's time, do it with your own data to at least help move the conversation forward.

It is fair to discuss a published analysis. Saying that only people who bring their own month of equivalent analysis (which conveniently would take another month to produce) are allowed to critique it is just a cheap trick to shut people down.

If you post big claims, they are open to analysis and review by others

show 1 reply