logoalt Hacker News

Web Search API

453 points • by tosh • today at 10:47 AM • 208 comments • view on HN

Comments

simonw • today at 2:29 PM

My number one question about search APIs is always if they allow you to store and resyndicate results you get from them.

If I'm running an agent system but I'm not allowed to store the responses - or provide a "share transcript" button - that's a pretty significant limitation.

The answer to that question is inevitably buried deep in the terms. Here's the relevant section I found for Ceramic, in their list of things you can't do:

> (n) collect, aggregate, store, or compile Output, including search results, relevance scores, or rankings, for the purpose of creating or contributing to any database, dataset, index, or corpus, whether or not such database, dataset, index, or corpus is used for a purpose that competes with Ceramic; (o) resell, syndicate, or otherwise make Output available to any third party on a standalone basis or as a separately accessible component of another product or service; provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query, and is not independently accessible, extractable, or downloadable by end users or third parties; or (p) retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use, unless expressly permitted in an applicable Order Form.

https://www.ceramic.ai/terms-of-service

Am I alone in caring about this?

➕ show 5 replies
iphonecorridor • today at 12:06 PM

For those developers out there, the best is still Gemini Flash Lite 2.5 believe it or not. It gives you 1000 google searches per day for free. Compare to Flash Lite 3.x which is 5k PER MONTH and then a few pennies PER SEARCH. Nuts. Didn’t realize search was so expensive.

Perhaps realizing all of this, Google hasn’t yet deprecated 2.5, bit limits access to it to “those who have used it before.”

It’s really really good for low cost search!

➕ show 7 replies
binarymax • today at 11:32 AM

Why not use those providers directly? Does Cloudflare need to be in the middle of everything?

➕ show 8 replies
qznc • today at 12:21 PM

My coding agent uses the hister cli, i.e. a local index. That often requires me to seed it manually as a downside. The upside is that it caches website contents via browser plugin, which is a nice workaround for bot blocking.

Thanks asciimoo for https://github.com/asciimoo/hister

➕ show 4 replies
jasonjmcghee • today at 2:02 PM

> All three support Zero Data Retention for requests made through Cloudflare

And then on the providers page:

    Property              Value
    provider              exa
    Zero Data Retention   No
➕ show 1 reply
karmakaze • today at 2:07 PM

I was just looking into these as DeepSeek Harness w/ Qwen3.8-27B relies heavily on search. I was going to go with Serper.dev[0] $1 per 1000 (or lower in quantity).

The providers[1] behind this Web Search API have very different rates:

    Ceramic.ai: $0.25 per 1,000 requests
    Linkup:     $5.00 per 1,000 requests
    Exa:        $7.00 per 1,000 requests
[0] https://serper.dev/

[1] https://developers.cloudflare.com/web-search/providers/

leflob • today at 9:52 PM

sorry this might be a dumb question but I am not really clear if they have a pricing structure and how much that is. I couldnt find any pricing directly linked to the Web Search API but then i see references that you use 'AI Gateway credits', but I also couldnt find pricing or free limits for those ones as well. Can somebody cue me in?

Herz • today at 12:06 PM

How does Cloudflare manage to hit the HN front page almost daily? Don't get me wrong, they build cool stuff, but the frequency is wild.

➕ show 4 replies
fnordsensei • today at 1:11 PM

I've been quite satisfied with Kagi[1]'s API.

1: https://kagi.com/api/docs/openapi

➕ show 3 replies
sreekanth850 • today at 1:07 PM

Create bot detection and bot protection, then sell crawlers. Is this the peak of hypocrisy?

➕ show 3 replies
jwr • today at 6:30 PM

Many people have outsourced the decision on who can access their websites to CloudFlare ("bot protection"), which incidentally makes these websites harder to access by bots working for humans.

Now there is an official paid search API, and I'm guessing the certified providers will be allowed through the Cloudflare "bot protection"?

This is very worrying.

freakynit • today at 12:14 PM

Tried one query on ceramic.ai (the default provider for cloudflare web search api): "qwen-3.8 flash next and rtx 5090 best inference setup" ... 0 results ... same query on google and ddg both yield proper results.

Then shortened the query to just "qwen-3.8 flash next" ... results came.. all unrelated. In fact, these were almost all paper links .... no relation to actual search term.

And I had thought that I finally had found a cheaper search alternative.

➕ show 2 replies
yellow_lead • today at 5:49 PM

Reselling APIs seems lazy to me. I wonder if CF plans to make their own provider. That's what I had assumed when I read the title.

➕ show 1 reply
hrideshmg • today at 7:01 PM

Surprised no one in this thread has mentioned running a self hosted search API.

I personally used to use Firecrawl's paid credits (got a bunch of em for free at an event) before I realized that they allow you to self-host your own instance (albeit missing some features I never use anyways).

It's been working really well for my agents, I even hosted a small observability tool that proxies the requests so I can see how many are failing and the percentages are always below 2%.

➕ show 1 reply
laumars • today at 5:52 PM

Lately I've been using SearXNG for personal models. It's free and seems ok thus far.

https://docs.searxng.org/

tom1337 • today at 11:32 AM

I wonder if the three search engines get access to cloudflare protected sites without any captcha or bot interventions

➕ show 1 reply
EcommerceFlow • today at 6:39 PM

Spent a few months building a product scraper using a mad mash up of various LLMs, OCR, etc. The pricing for their providers is 3x-8x higher than something like Luna 5.6 w/ Web Search. Not sure what their differentiator is, unless they just wanted to launch something.

0fflineuser • today at 2:17 PM

I am pretty sure exa specifically say it trains on your data in it's privacy policy, so how can it be ZDR ?

I remember as I was looking at the available web tools for hermes agent not to long ago and looked through the keyless web providers privacy policies, which exa is one of them.

➕ show 2 replies
8bite • today at 1:44 PM

The pricing is so different between these:

ceramic.ai - $0.25 per 1,000 requests

Exa - $7.00 per 1,000 requests

Linkup - $5.00 per 1,000 requests

Does anyone have insights on the quality differences? Web search API pricing for AI agent usecases has always felt so expensive for what it is, but I have no grounding on the economics of running a web index.

EDIT: formatting

➕ show 4 replies
solaire_oa • today at 7:27 PM

Anthropic plausibly uses Brave Search... and Brave search maintains its own index. Makes sense: cheaper search API, leveraging non-Google, etc.

Here we are, one layer of indirection more: Ceramic, Exa, Linkup. Who knows what they use. If you told me that those 3 build and maintain their own index, I'd first question whether that was true, and if it is, I would question whether it was any good (relative to Google/Bing/Brave).

So what is CF providing here? Maybe some free credits to entice us to use their router? No, not that either ("billed to your AI Gateway credits"). Maybe a comparison of which agent search yields the best results? Nope.

It's a crappy proxy- probably less efficient and more volatile than hitting the agent API directly.

This is only if I understand the product correctly (which I admittedly skimmed) due to the sheer number of screeching vibey nothingburgers coming out of CF over the past month.

➕ show 1 reply
blurbleblurble • today at 7:26 PM

At first I thought this was a new web standard and was intrigued to hear what they'd come up with, sad to find otherwise

chews • today at 10:01 PM

I love that businesses that would've been considered too risky to get in are all the rage if they feed the LLM demon. Scraping other's results? NO PROBLEM! They carry so much traffic they can literally just syndicate their network pipeline and find a new bullets for the money gun.

starcast2026 • today at 6:04 PM

I am trying to understand the value.. This is for customers who have their agents already on CF? Improved latency & same eco-system etc., Right? Because others can always use Google Search APIs

➕ show 1 reply
blakeashleyjr • today at 7:02 PM

When I saw this, I assumed CF was going to offer an API to access the pages they otherwise protect.

NO SCRAPERS (except ours) -> $$$$$$$$$$$$$$$$

hmartin • today at 3:21 PM

I've been working on a TypeScript package to provide a unified search API across these providers:

https://github.com/hbmartin/agent-web-search

So this gives a unified search experience without adding another cloud hop and dependency.

saltysalt • today at 2:47 PM

Wow I guess I am in the minority of folks building a search engine for humans now, this is a wild business model but best of luck to the 3 search index providers sitting behind this proxy, I hope it's worth their while financially speaking. Building an index is hard and expensive (I know).

anon373839 • today at 12:09 PM

> All three support Zero Data Retention for requests made through Cloudflare

But does CloudFlare itself commit to zero data retention? If not, this isn’t too meaningful.

duncangh • today at 5:06 PM

Cloudflare has been shipping more than FedEx lately. Would love to learn more about how they are going about this from a strategy, planning and execution standpoint.

ronfriedhaber • today at 1:09 PM

Interesting to see how this can be compared with Exa, Alas, Cloudflare really is shipping many great orthogonal products recently.

➕ show 1 reply
Oras • today at 11:59 AM

Weird choice by CloudFlare, would been great if they have shared why it was created.

I use CloudFlare developer platform and quite happy with tools, but I didn’t use the gateway API and always used OpenRouter which does support web search.

I can see it useful for those who didn’t do any integrations or like to keep logs at one place, but did customers actually ask for this?

innocent_name • today at 5:06 PM

Of course Firecrawl isn't "Verified bot". Their customers were responsible for 95% of my traffic bill overcharge.

daft_pink • today at 4:38 PM

Wow, will they offers some sort of reduce a web page into a markdown file api as welL so we can get full web pages at reduced token sizes?

➕ show 2 replies
maelito • today at 4:50 PM

Same as https://staan.ai, the European index.

vscarpenter • today at 12:20 PM

I guess I'm not understanding the value here - to compete with Google and the likes, the scale, cost and complexity would be huge. Appreciate new entrants in an existing field but not seeing this one.

➕ show 2 replies
ByeByeSpace • today at 5:10 PM

I run a small side project on Cloudflare Workers, so having search available right from a Worker without adding another vendor is appealing. Curious how the pricing compares to Brave's search API.

babelfish • today at 4:49 PM

What value does Cloudflare provide over using the linked providers directly...?

corentin88 • today at 1:22 PM

Funny how search was a graveyard for startups for almost two decades. And since ChatGPT releases (or so) it’s a trending place again.

➕ show 1 reply
asjq178 • today at 4:33 PM

Cloudflare protects against bots, Cloudflare sells out its customers to AI scrapers.

MITM service, Internet gatekeeper and robber baron.

➕ show 1 reply
0xbadcafebee • today at 4:24 PM

SearXNG works pretty well for my personal agents. FYI it's a free search gateway you can host locally, and there are many public instances. It's like the old days when many different people provided the same free service for all.

okokwhatever • today at 12:31 PM

Who thought 5 years ago that searching online would have a cost...

➕ show 2 replies
kobieps • today at 5:10 PM

so this is like a competitor to https://parallel.ai/ ?

➕ show 1 reply
Jeeetendra • today at 2:44 PM

the zero-retention promise from the search providers is useful, but the requests also show up in gateway logs. can you keep the billing data without storing the actual search queries?

timpera • today at 11:51 AM

Interesting to see that they didn't include the Brave Search API, which is really great and imo a better experience than Exa.

The absence of the Perplexity Search API is to be expected though, knowing how much these two companies despise each other.

➕ show 1 reply
1vuio0pswjnm7 • today at 4:45 PM

Explore Cloudflare's services without the web page bloat:

https://developers.cloudflare.com/llms.txt

As a textmode command line and text-only browser user this textfile is faster for me to use that the usual Silicon Valley style web pages

Not quite as good as sitemap-0.xml but it's nice to have this in addition

Nan0pk • today at 5:38 PM

why are they selling search as a paid feature...

measurablefunc • today at 5:21 PM

Very cool.

zergrush • today at 5:19 PM

is this serpapi ?

🔗 View 9 more comments