logoalt Hacker News

Show HN: Let your AI agents paint big arrows, boxes and text on your screen

354 points • by franze • today at 11:03 AM • 150 comments • view on HN

Comments

sicktriple • today at 4:29 PM

Man, just when it seemed like we had it all. Computers were cheap, efficient and powerful. Somehow we figured out a way to accomplish tasks we already had solved except now it's 1000x more expensive, requires the combined electricity of the entire world, is reliant on someone else's rented compute, and now I need a robot to tell me what button to press. What a time to be alive.

➕ show 7 replies
hn8726 • today at 12:21 PM

I tried to read the "Does it need Screen Recording or Accessibility?" part, but it's slopped to the point I have no clue what it's trying to say. But if it can draw on top of permission prompts, what's stopping it from drawing box that hides the "decline" button and changing the "approve" button copy?

➕ show 5 replies
internet101010 • today at 8:51 PM

The worst trend in UX in the last decade is the endless "Got it!" popups and feature notifications that distract the user from what they were trying to do.

I don't know why anyone would ever willingly want this.

tangotaylor • today at 2:41 PM

"It is an arrow, so we spent an unreasonable amount of time on how it looks."

Brilliant. This is exactly the kind of content I seek when I visit Hacker News.

Truly art.

➕ show 1 reply
usrbinbash • today at 1:14 PM

SO the point of this is ... what exactly?

A big arrow to an interface element which ... has a label that explains what it does?

So...a label for a label?

➕ show 5 replies
arshxyz • today at 11:55 AM

The README is geared towards technical people (complete with the HN screenshot) but when I see a tool like this all I can think of is how helpful this would be for my mom when I'm trying to tell her how to download and print a document over the phone

➕ show 2 replies
kogus • today at 5:46 PM

My first reaction was similar to the reaction I'd have if you told me that cockroaches had learned to unlock doors and stand on their hind legs. But then I thought about accessibility, and the ways this could be used to help technologically illiterate or disabled people, and I thought again.

Do you remember when PCs used to come with a completely soup-to-nuts tutorial that would talk to you like you had never seen a PC before? Things like this: https://www.youtube.com/watch?v=3ScS4OYDfHE

This kind of baked-in interactivity could really help in a training or disability context.

isoprophlex • today at 11:43 AM

Literally unusable as it is. Some minimal extra features this would need:

- rainbow dripping arrows

- angrily pointing arrows

- flame-surrounded text boxes with particle effects

- the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings

EDIT: ayy lmao https://github.com/franzenzenhofer/big-arrow-on-the-screen/p...

➕ show 3 replies
lbreakjai • today at 12:15 PM

I would pay good money for something like this on iPad. It wouldn't even need to be agent-driven, just a big "I want to make a bank transfer" button, that would launch the correct app and guide through the interface.

That would be a godsent for those of us with aging parents.

➕ show 1 reply
alansaber • today at 4:16 PM

OP has inspired me to write up my tedious thoughts on agent GUIs if of interest https://news.ycombinator.com/item?id=50022688

cyberjunkie • today at 12:39 PM

I'm just as impressed by this as any other LLM-generated project.

priyashunt • today at 4:26 PM

Very VERY USABLE FOR old PEOPLE. I would pay good money for something like this on iPad!

vessenes • today at 11:34 AM

Interesting. When I read the headline I imagined this would be a sort of thinking trace booster -- letting the agent focus its own attention on different parts of the screen. But this is cool in a different way. I bet agentic harnesses would find it useful for communicating with other agents / themselves as well.

alansaber • today at 12:06 PM

This might be goofy, but it underscores that there's potential for more visual agent UIUX than reading off a sidebar/opening modals.

➕ show 1 reply
melvinroest • today at 12:12 PM

My message to the world is that LLMs should be able to point anything they see in the application they're in or even the whole computer (if you give it that kind of access).

For web apps, WebMCP is a way where you can make a frontend way more discoverable to an LLM rather than that it is going to read all your code. I sometimes create this in my apps at home and sometimes in my apps at work and it makes LLM assistants way more usable. It also helps if they can highlight certain things of a visual.

We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?

FinnLobsien • today at 11:42 AM

This could be great for documentation. Screenshots in docs are frequently useless because they show me a screen and say "click X" where I still have to search X visually. And I could just to dhat in the other tab I have open.

➕ show 1 reply
alexpotato • today at 2:14 PM

> Arrows have existed since roughly the Paleolithic.

Whenever I use web tools that don't have "deep linking" [0], I love to throw out this quote:

"Have we thought about using hyperlinks? They are these SUPER useful things that were invented in Switzerland back in the 1990s."

0 - https://en.wikipedia.org/wiki/Deep_linking

swframe2 • today at 4:11 PM

I want this for the visualization of very difficult to solve application evolution. This is for problems which have no known solution and are too complicated for an agent to figure out on its own. My current problem: can alphafold and related tools figure out the function of a gene that so far is unknown. Think of the monte carlo tree search in the alphago explanation videos. What if there was visualization that showed how the policy and value models worked so you can spot their flaws. I want the agent to build an attempt at a solution, then build a visualization of it, then run the solution and show me what is it up to. I want to it pause and explain its state so I can see exactly where and why the solution fails.

satyanash • today at 11:22 AM

Am I missing something here?

What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?

If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.

➕ show 2 replies
dr_kiszonka • today at 4:06 PM

I take this opportunity to shame GitHub for completely ignoring the mobile experience in their own Android app. The project's README is pages upon pages of largely blank space. In general, the app does not render mermaid diagrams and does not allow for zooming in, so smaller pictures are unusable. Even if you access pictures directly in a repo's source, you still can't zoom in. iPython Notebooks, which are extremely common in data science, are an "unsupported file type" and are not rendered.

(OP, nice project! Sorry for my rant.)

➕ show 1 reply
ElijahLynn • today at 8:03 PM

This is going to be really useful actually, nice work! Can't wait for it to be on Linux too!

m-s-y • today at 1:07 PM

Genuine question…

Why does this need to be a skill? Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows” and you get the same thing without burning tens of thousands of tokens in context.

➕ show 2 replies
TekMol • today at 11:31 AM

Swift, Shell, Python and Objective-C

Does one need 4 programming languages to draw something on a mac?

➕ show 1 reply
iandanforth • today at 4:12 PM

Pointing is surprisingly overlooked in a lot of tools and I look forward to using this. It reminds me of what I thought was a killer feature on AnyBots telepresense robots, they had a laser you could use to point at things while piloting the robot.

ghm2180 • today at 12:52 PM

Man, Ive lost count of How many times have I had to repeat this over the phone to my parents; The repo has all the punch lines in the README

> "It's this window, not that one." You have 14 Chrome windows. The agent knows which one it means: --window "Google Chrome:Pull request". It even picks the right tab: --app "Google Chrome:Pull request".

> Remote help. "No, the other gear icon." Point at it instead of describing it.

xyzsparetimexyz • today at 11:44 AM

Seems like a pretyu useful way to help infants use desktop computers

➕ show 1 reply
ceroxylon • today at 5:31 PM

This is like the alerts in new cars that remind you to check for other passengers when you exit... helpful, but a worrying sign of the times.

LoneRanger1024 • today at 3:28 PM

I think this makes a lot of sense for tutorials, but permission prompts need different treatment. Besides pointing to a button, the assistant should explain what clicking it will do, so the user can make their own decision.

tilemarch • today at 2:04 PM

An arrow on the screen solves “where do I click”; it doesn’t solve “should this happen”

To me this feels like it takes away from what the human is supposed to do (read, understand the consequences of the action, then.. consent or abort)

There is a reason your AI Agent won't automate these clicks for you

user- • today at 5:10 PM

Seems great for scammers targeting old people tbh

code_duck • today at 5:37 PM

This is definitely not something I want, ever. Maybe someone could be helped by it I guess.

flr03 • today at 12:35 PM

That would have been handy 20 years ago to point to that one valid 'Download' button.

➕ show 1 reply
pimlottc • today at 1:23 PM

Even the example arrows in the first screenshot are wonky and unnaturally weird.

antonyragleap • today at 3:45 PM

Clever approach for agent debugging. Visual cues > logs when multiple agents run. How do you handle overlapping annotations?

ForHackernews • today at 11:20 AM

"and they keep hitting the same wall, the part that only a human may do"

Grim. If you're just there to click sudo buttons for the bot, you might as well give it root access and be done.

➕ show 2 replies
ex-aws-dude • today at 12:14 PM

If you can’t even take the time to understand what you’re clicking why even go through the formality of “approving”

harrouet • today at 1:32 PM

There is no limit to burning tokens :)

➕ show 1 reply
amelius • today at 1:12 PM

Because AI can paint pelicans on bicycles quite well, and not arrows?

lapestenoire • today at 11:27 AM

I love it.

Retr0id • today at 12:35 PM

Reminds me of something from Idiocracy (2006)

wartywhoa23 • today at 6:11 PM

Next up: Spare yourself from formulating any prompts, let our SI agent control our thin client by asking you leading questions that you answer by drooling for yes and picking the nose for no!

okasaki • today at 2:56 PM

Look at what Windows and Mac users need to mimic a fraction of our power

rrr_oh_man • today at 5:50 PM

HA! Franz!

I remember taking your SEO course many moons ago.

erickhill • today at 3:49 PM

I don't know why exactly but something about the arrows feel repulsive and oddly gross.

➕ show 1 reply
emsign • today at 5:16 PM

Reminds me of Idiocracy. Are we idiots yet? This is so that AI agents can "guide" a human, directing the meat bot more easily. Telling the meat bot where to click and what to do. The entire "what to think" step isn't even necessary when there is no thought. Because thoughts and decisions are annoying. It's astonishing how much people are already "algorithmically" steered so they don't have to turn on their rational thinking.

➕ show 1 reply
LetsGetTechnicl • today at 6:22 PM

Really groundbreaking stuff guys, another $100 billion to you

intended • today at 2:17 PM

> big-arrow-on-the-screen (bigarrow): one small macOS CLI and an agent skill. Click-through, never steals the focus, gone by itself. MIT.

I absolutely loathe this phrasing now. I don’t even know what part of it is good or notable.

DonHopkins • today at 11:57 AM

Ha ha, I love it! I wrote a pointing hand annotation overlay in PostScript in 1989 for NeWS and the PSIBER Space Deck's Pseudo Scientific Visializer:

  %  @(#)handy.ps
  %
  %  Handy Pointer
  %  Copyright (C) 1989.
  %  By Don Hopkins. ([email protected])
  %  All rights reserved.
https://donhopkins.com/home/archive/psiber/cyber/pointer.ps

PSIBER Space Deck and Pseudo Scientific Visualizer Demo:

https://youtu.be/_fqCeuue5Ac?t=213

The Shape of PSIBER Space: PostScript Interactive Bug Eradication Routines — October 1989:

https://medium.com/@donhopkins/the-shape-of-psiber-space-oct...

➕ show 1 reply
dihinbutt • today at 3:21 PM

I love this

npodbielski • today at 2:07 PM

seems useful for most of the folks that just want to do stuff without going through (over)complicated UIs...

But I feel dumber just by looking at it. If this is how this will look like in 10 years, why to make desktop at all. Just connect mic and speaker to your PC or talk to the phone: 'I need new pair of socks. Order 10 for me. In your favourige color it will be 21.37. Should I charge your credit card?'.

It is not like most of the people enjoy computers. I am pretty sure they do not. They just need them to operate systems they need: government websites, banks, maps, restaurant menus etc. If some agent will do that for them, why bother looking at screen at all? Rich people have it with their own personal assistants.

🔗 View 11 more comments