logoalt Hacker News

simonw • today at 2:29 PM • 6 replies • view on HN

My number one question about search APIs is always if they allow you to store and resyndicate results you get from them.

If I'm running an agent system but I'm not allowed to store the responses - or provide a "share transcript" button - that's a pretty significant limitation.

The answer to that question is inevitably buried deep in the terms. Here's the relevant section I found for Ceramic, in their list of things you can't do:

> (n) collect, aggregate, store, or compile Output, including search results, relevance scores, or rankings, for the purpose of creating or contributing to any database, dataset, index, or corpus, whether or not such database, dataset, index, or corpus is used for a purpose that competes with Ceramic; (o) resell, syndicate, or otherwise make Output available to any third party on a standalone basis or as a separately accessible component of another product or service; provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query, and is not independently accessible, extractable, or downloadable by end users or third parties; or (p) retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use, unless expressly permitted in an applicable Order Form.

https://www.ceramic.ai/terms-of-service

Am I alone in caring about this?


Replies

ChuckMcM • today at 10:57 PM

No, and it's funny that the "data" they scrape for "free"[1] they forbid you from reselling. It is entirely unclear to me if this is even enforceable.I mean who are the parties in this transaction? Two programs? If you're responsible for what your programs do why aren't AI folks going to jail for violating CFAAA? Hmmm? It is all kind of screwed up.

When we were running Blekko we had a lot of people who were attempting to crawl WordPress sites for plugins that had unpatched vulnerabilities or shopping cart packages with the same. I'm really curious though about how Cloudflare is monetizing this.

[1] We all know its not free to support the crawling traffic of a web scraper.

infogulch • today at 4:12 PM

It seemed like this part gives you the exception you wanted:

> provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query ...

but it continues:

> ... and is not independently accessible, extractable, or downloadable by end users or third parties

How can you prevent end users from extracting it if its visible? Why even have the exception if you just throw it out with an impossible to meet restriction like this?

➕ show 2 replies
cj • today at 3:51 PM

My general stance on things like this is to think about the intent -- why does the company have that in their TOS. Use that as a proxy for assessing the likelihood of the company enforcing the terms against you.

➕ show 1 reply
sanderjd • today at 2:38 PM

You are not alone, I agree that this is a frustrating limitation.

phoghed • today at 5:21 PM

It’s a shit tier web scraping startup, just violate their terms, who cares.

➕ show 1 reply