logoalt Hacker News

criley2today at 11:34 AM2 repliesview on HN

[flagged]


Replies

wizzwizz4today at 11:56 AM

> Try some context engineering. Try customizing your harness. Try having the harness improve itself. These are AI 101 lessons

If AI is really that complicated, it sounds like it would be easier to write an ordinary computer program to aggregate job boards.

show 2 replies
tcp_handshakertoday at 11:41 AM

Its the tools. Sell those RSUs while they last.

Most of these are 2026....

Frontier LLMs Still Struggle with Simple Reasoning Tasks - https://arxiv.org/abs/2507.07313

General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks - https://arxiv.org/abs/2604.11778

LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening - https://arxiv.org/abs/2605.19597

LogicGraph: Benchmarking Multi-Path Logical Reasoning via Neuro-Symbolic Generation and Verification - https://arxiv.org/abs/2602.21044

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models - https://arxiv.org/abs/2607.08317

Vision-Language Models Lag Human Performance on Physical Dynamics and Intent Reasoning - https://arxiv.org/abs/2601.01547

Do Vision-Language Models Understand 3D Scenes or Just Catalogue Objects? - https://arxiv.org/abs/2605.20448

The Reversal Curse: LLMs Trained on “A is B” Fail to Learn “B is A” - https://arxiv.org/abs/2309.12288

Large Language Model Reasoning Failures - https://arxiv.org/abs/2602.06176

show 1 reply