logoalt Hacker News

Usenet rewind archive search engine

126 pointsby cstadler1869yesterday at 4:19 AM38 commentsview on HN

Comments

vincyesterday at 3:28 PM

You can visit only 25 pages on the website with a free account, that's a bit too limiting. I'd suggest limiting the number of searches maybe but not the pagination over the results.

I got curious again to find my old posts and figured I could download and grep the mbox files from Internet Archive instead.

msephtonyesterday at 12:19 PM

Just found some cool old stuff thanks to this:

1996-01-17: first post! a reply to a thread in alt.music.bjork

1996-03-11: first mention of my website/homepage

No links to avoid my own embarrassment.

show 2 replies
rfarley04yesterday at 4:38 AM

Oh my god this is a lifesaver. Just last week I was using Google Groups to search for old Usenet posts about the Commodore 64 and the Mimic Spartan, and I was getting "too many requests" errors 4 out of 5 times clicking a result. Happened on VPN+incognito, and in a separate browser. Just searched for what I was looking for in Rewind and immediately found the posts and even some I never saw in Google Groups results. So glad all of this is archived in multiple, non Google places

show 1 reply
kingroloyesterday at 11:09 AM

Ah this has reminded me of dejanews.com the usenet archive. Before the days of Stack Overflow that was absolute gold for looking up an error message or code issue.

If I remember rightly Google bought the data when they shutdown and incorporated it into Google Groups.

show 1 reply
matherialyesterday at 6:33 AM

Selfishly, I hate these things. No one in the 1980s and early 1990s expected the internet to become what it did. I discovered it as a dumb teenager and posted some frankly idiotic stuff. I also put my real address and phone number in the signature, because that's how you rolled back then, the entire internet was like 10 nerds.

Now, 35 years later, a lot of other stuff has mercifully decayed, but people keep bootstrapping these Usenet archives as pet projects every year, and there's never an option to opt out. Today, agents make it easier than ever to build a dossier on anyone you don't like, so this data is bound to cause some harm. And I guess when I die, the best-preserved memory of me that my grandchildren and grand-grandchildren will be able to pull up will be snarky posts from a 15-year-old on some unix-related discussion group.

Yeah, I know there's a bunch of posts there that are of genuine historical interest. But I'm just not a fan and I don't mind burning karma to speak my mind.

show 5 replies
fros1yyesterday at 11:56 AM

Logistically, these collections usually exclude alt.binaries traffic entirely. I often wonder if there are enough obsessive collectors out there to piece together an archive from ancient tapes and hard drives.

sparrowidleyesterday at 9:51 AM

Grabbed the UTZOO tapes off archive.org a while back, ~2GB compressed for 1981-1991, and recoll indexed the whole thing in an afternoon.

show 1 reply
hn1rig3rakyesterday at 10:05 AM

If you go local, notmuch beat recoll for me on the mbox split, ~2M messages indexed in about 40 min on a cheap SSD.

ChrisArchitectyesterday at 6:35 AM

A similar one from earlier in the year:

Usenet Archives

https://news.ycombinator.com/item?id=47655905

show 1 reply
ggmyesterday at 5:36 AM

The most prolific self-references (to myself that is) I found were UUCP maps updates because I ran a node which was in the centrally coordinated catalog (iirc it drove honey-danber to convert a!b!c paths to user@c so you didn't have to hand-route)

The second-most prolific self references were blovating across all and any groups I felt like. Nothing has changed.

show 1 reply
retracyesterday at 5:26 PM

Very cool. I've been working on something similar. My cutoff date is 1994 as Usenet changes substantially in nature then.

May I ask how you collected the data? Usenet is not fully archived anywhere. No one has a complete set.

I am aware of:

* the NewNews disc series on CD from '91 - '93 which has a substantial text feed for the big groups during that era (most of these are on IA.org, but some are not!)

* The Dejanews archives via Google up to c. 2013. Also up on IA.org These are lacking large amounts of early posts, and are highly spotty before about 1994 (when Dejanews started).

* The Utzoo tapes over at https://archive.org/details/utzoo-wiseman-usenet-archive cover a part of '81 - '91 but only from one host. "Far parts" of Usenet over in Japan or Europe are barely covered.

* Some groups preserve per-newsgroup histories scattered all over and tending to drop offline over the years

* fj.* is well-archived but it's Japanese

The chart speaks for itself I think as to coverage: https://i.imgur.com/WCQm4od.png

I got so desperate as to go scanning zip and tar files on old ftp sites for anything with a Usenet header. That scored a few hundred thousand articles mostly from the 90s.

The big spike in the early 90s is from those CDs archives in the months where I had access. It gives an idea for how limited those Giganews/Google and other surviving archives are.

I figure less than 50% early text Usenet survives, unless someone has vast hordes of data that they are not disclosing.

As for the Utzoo archives - they're legally encumbered and some of the posters sue anyone who hosts them. (They're not available from the IA - taken down.)

You have a takedown page but the legal and copyright issues were one of the main back-burner items keeping me from making my archive live. Some of the more eccentric Usenet personalities are still alive.

Same with privacy issues. I notice you have full text search. I was reluctant to do that for privacy reasons.

I discovered very early on how to index by email but that felt quite invasive. I took that out, and built a graph of all the X-refs: which turned out to be more interesting (walking the discussion trees) and somewhat more privacy preserving in that you can explore but not search.

I was also thinking of an AI-based content warning filter, because I have a "random post" button and it is sometimes wild what that lands on.

alfiedotwtfyesterday at 2:04 PM

Where is the source of all these Usenet services? I thought people just stopped mirroring when Google groups got so prolific? I wonder how big all of this is minus the *.binaries

doublerabbityesterday at 11:32 AM

It's nice that usenet still exists today, although if not mainly for warez.

It's sad that it's now powered by only four usenet backbones.

TrueRideryesterday at 8:11 PM

[dead]

ember_ilandsyesterday at 3:44 PM

[flagged]