logoalt Hacker News

I accidentally turned LLM memory into program analysis

71 pointsby matt_dyesterday at 11:27 PM12 commentsview on HN

Comments

iamflimflam1today at 4:49 AM

This really matches up to my experience on long research projects with Claude.

It’s very hard to remove information - Claude has a habit of recording things all over the place and will happily treat things as facts even after they’ve been disproved.

What is currently true can get easily contaminated with old “facts”.

sim04fultoday at 4:42 AM

I reached a similar conclusion: LLMs should only really sit at the terminals of request fulfilment.

1. User request understanding: natural language -> a more rigorous representation, in my case Datalog.

2. Result interpretation: facts and derived facts -> natural language.

Between those terminals, the work should be mechanical reasoning over some ontology or formal knowledge structure.

That connects to another principle I've been thinking about, which I call Weathering: useful reasoning should leave durable residue in the system. If an LLM has already had to infer a relation, mapping, rule, or abstraction, the next similar request shouldn't pay the full cost of discovering it again.

With continued use, the system should therefore require less and less probabilistic intelligence for recurring work. Put another way, there should be a declining marginal cost of cognition: the products of intelligence harden into structure that can subsequently be reused and evaluated mechanically.

trinsic2today at 1:15 AM

Something of this capacity would be useful in investigating obscure hardware failures in the logs that I couldn't confirm because the problem was not being observed while the device was in my shop. the problem was surfacing in another location probably due to some set of circumstances in the software that I could recreate, or some particular peripherals that were attached.

I ran into the very same problem of the LLM forgetting that we ruled out a conclusion that was verified not to be the cause as it came up further in the conversation history while I was exploring possibilities.

I had to keep reminding we ruled out that conclusion prior.. I just carried on with having the LLM capture some of the supporting sources of other people experiencing the same problem and kept having to refine those sources because it was focused only on summaries, but eventually i got the sources to a point where they were good enough hypothesis that we could formulate a better conclusion on what the potential cause was.

keedatoday at 1:41 AM

Very cool. I recall an HN submission (which I can't find offhand unfortunately) that did something similar -- it used an LLM to decompose articles into a set of statements which were used to construct an entity-relationship graph of facts and events. It then queried that using conventional graph query methods, much like DataLog / Lemmalog is doing here. I remember it was particularly effective at answering timeline-based queries that LLMs (back then) sucked at.

(See also Cyc: https://en.wikipedia.org/wiki/Cyc)

I think approaches like this are going to be (or maybe already are?) the basis of effective grounding of LLM responses in authoritative data sources. It should be possible to pinpoint any error to an incorrect traversal or an incorrect "fact." This would work best for concrete, unambiguous facts, however; fuzzy, ambiguous or opinion-based information will probably remain the purview of LLMs.

esttoday at 3:49 AM

Very cool article. I had a similar idea where "fact checking" should be real programs for logic correctness.

But IRL it's too vague. The exploit hunting is a better use case.

vatsachaktoday at 3:18 AM

Eventually lambda prolog will rise again

linguaeyesterday at 11:48 PM

This summer I’ve been investigating agentic coding with local LLMs, and while I’m far from an expert, one thought that has been on my mind is leveraging techniques from “old-school” AI such as heuristic search to guide agents when it comes to planning. The use of Datalog in this article resonates with me, since logic programming was a major part of old-fashioned symbolic AI. I’m very curious about this combination of “old-school” AI and LLMs.

ande-mnoctoday at 3:23 AM

Ctrl-F “prolog”: 0 result. :-/

show 1 reply
fizxtoday at 1:05 AM

Is this sort of re-inventing Graph RAG from another angle, or does it feel novel?

show 1 reply
langstoday at 3:29 AM

[dead]

fenestellatoday at 2:24 AM

[flagged]