GoldieBench Blog · 9 min read
Hark AI Agent: What It Is, The Model Behind It, And How Those Brains Score On GoldieBench
The Hark AI agent doesn't officially name its model. Here's what's reported, how those brains score on GoldieBench, and how to set Hark up free.

The Hark AI agent is Hark Pro, a free personal agent from Hark Labs that connects to your email and apps and does real tasks on its own cloud computer.
It launched on 6 October 2026, and I signed up on the free plan the same day to test it on camera.
The first question builders ask me is simple: which model is actually running it?
I run GoldieBench, where 40 models have now produced 1,609 one-shot demos across 50 tasks, so I can at least tell you how the likely brains score.
I'll be plain up front, though.
Hark is an app rather than a model, so it isn't on the board, and Hark has not officially said which model powers its main assistant.
So this post explains what Hark is and how to use it, then lays out what is known about its models and how those models score on GoldieBench.
What the Hark AI agent is
Hark describes Hark Pro as "the operating system for your life".
It chats like a messaging app, remembers your preferences, and uses its own cloud computer, called Handoff, to book, buy, plan and pay.
Hark says Handoff can juggle up to six browsers at the same time and log in to websites on your behalf.
Hark Labs was founded by Brett Adcock, who also founded Figure and Archer Aviation, and TechCrunch reported a $700 million Series A for Hark in May 2026.
When I opened it, it felt like the Apple of AI agents, because the design is that clean.
| Plan | Price | Usage |
|---|---|---|
| Hark Pro | Free | Every feature with base usage |
| Hark Pro² | $20 a month | 2x more usage |
| Hark Pro³ | $100 a month | 10x more usage |
Those plan names and prices come from Hark's own launch post, which also says every feature will always be free.
It runs on the web at hark.com and in the Hark Pro apps on iOS and Android, and Hark's terms say you must be 18 or over.
How to set up the Hark AI agent
Setup took me a few minutes.
First, sign up at hark.com or install the Hark Pro app.
Second, connect your email, because Hark recommends doing that first, and I connected my Gmail.
Third, check the details Hark shows you, such as your location.
Fourth, answer its opening question, which for me was, "Hi Julian. What's one thing on your plate this week I could take off?"
Fifth, connect your other apps, such as Slack, GitHub, LinkedIn, Notion, Amazon and PayPal, and add your wallet only when you're ready.
Connected apps work directly inside Hark's own computer, which is easier than Grok Bot, where I often had to log in again through its computer.
Every Hark feature I tested
| Feature | What it does |
|---|---|
| Restaurant booking | I asked for an Italian restaurant next Wednesday, it asked for numbers and time, understood my typo'd "6pm, two people", and started booking until I said stop |
| Integrations | Gmail, Slack, GitHub, LinkedIn, Notion, Amazon and PayPal, plus Google and Outlook APIs and MCP |
| Wallet | Cards and passwords sit in the encrypted Secured by Hark vault |
| Home | Spaces, Panels and Action Buttons on one screen, and you can even pay bills from it |
| Panels | Mini apps you create by typing, such as a productivity panel with ClickUp or Asana, Shopify packages, calendar events or new Apple Podcasts episodes |
| Projects | Dedicated threads that store files for long-horizon work like an event |
| Proactive inbox | Reads your email and proposes action items for a yes or no |
| Side prompts and "show" | Prompts like "quiet this" and a show function that pulls items into the chat |
| Attachments and voice | Files and voice input are accepted, but there is no voice output |
| Scheduled tasks | One section that manages every recurring job you set in chat |
| Prep my day | Builds your itinerary from calendar and inbox |
| Trip planning | Finds camping spots, writes a packing list, says when to leave and books flights |
The "Prep my day" example I saw said, "Your inbox is clean.
Do you want me to pull in the sales numbers before your team sales meeting at 4 p.m.?"
Hark also says it's designing its own hardware, and launch coverage points to 2027 for the first devices.
Which model powers the Hark AI agent?
Here is everything that is actually known, split into what Hark says and what has been reported.
Hark says Handoff, its browser agent, is its own computer use model, which it post-trained itself.
Hark claims Handoff took the top spot on the Online-Mind2Web human evaluation leaderboard and beat GPT 5.4 by 8 points and Opus 4.8 by 2 points on average across its tests.
Those are Hark's own numbers, so treat them as vendor claims.
Hark has not named the model behind the main chat assistant.
TestingCatalog reported that Hark's configuration labels a model called Claude Opus 5.5 for the main assistant, task agents and website building.
The same report listed Gemini 3.1 Flash Image for images, plus Qwen3.6-27B, Gemini 3.5 Flash Lite and GPT-6 Astra for other jobs.
TestingCatalog itself warned that these labels do not confirm which model handled any real request.
So the honest answer is that Hark's main brain is undisclosed, and a Claude Opus 5.5 label is the strongest public clue.
How Hark's likely brains score on GoldieBench
GoldieBench gives every model the same one-shot prompt to build a game, a page, a simulation or a visual, then scores the rendered result from 0 to 10.
Here is how the models linked to Hark score on the live board today.
| Model | Link to Hark | GoldieBench avg | Tasks scored |
|---|---|---|---|
| Claude Opus 5.5 | Reported label for Hark's main assistant | 7.57 | 50 of 50 |
| Opus 4.8 | Hark compared Handoff against it | 7.51 | 47 of 47 |
| Gemini 3.6 Flash | The nearest Gemini Flash on the board, since 3.5 Flash Lite and 3.1 Flash Image are not benched | 7.08 | 50 of 50 |
| Qwen 3.7 | The nearest Qwen generalist on the board, since Qwen3.6-27B is not benched | 7.00 | 47 of 47 |
| Qwable 5 27B Coder | A 27B Qwen-family local coder for comparison | 7.14 | 50 of 50 |
GPT-6 Astra and GPT 5.4 are not on the board, so I won't guess their scores.
For context, the highest single model on the board is Claude Opus 5 at 8.27, with GPT-5.6 Sol at 8.16.
So if the Opus 5.5 label is right, Hark's main assistant sits in the solid upper-middle of the board rather than at the very top.
Claude Opus 5.5 rendered all 50 of its one-shot builds with zero console errors, and all 23 of its games passed the input playtest.
Its best page build, the web OS task, scored 8.7 and took a bronze medal.
That matters for Hark, because the same reported label covers website building.
What GoldieBench can and cannot tell you about Hark
GoldieBench measures how well a model builds working things in one shot.
It does not measure browsing, booking a table or reading your inbox.
Handoff is a separate computer use model, and the board does not test computer use at all.
So a strong bench score tells you the brain can reason and build, but it doesn't tell you Handoff will click the right button on a booking site.
That's why my restaurant test still matters more than any score for everyday use.
The best way to judge Hark is to give it one small real task and watch what it does.
How the other consumer agents' brains compare
Viewers keep asking how Hark stacks up against the other personal agents, so here are their brains on the board.
| Agent | Likely brain | GoldieBench avg |
|---|---|---|
| Hark Pro | Undisclosed, with a reported Claude Opus 5.5 label | 7.57 for Claude Opus 5.5 |
| Grok Bot | xAI's Grok models | 7.15 for Grok 4.7 and 8.09 for subscription Grok |
| Meta Muse | Meta's Muse Spark | 7.55 for Muse Spark 1.2 |
| ChatGPT Dots | GPT-6 Astra | Not on the board yet |
| OpenClaw | Whatever model you plug in | It depends on your choice |
The scores are close, and that's the real lesson.
The brains behind these consumer agents are all strong, so the differences you feel come from setup, design and how each agent handles your accounts.
On that front, Hark's onboarding was the easiest I've tried, and I called it a very simple, easy-to-use version of something like OpenClaw.
What a personal agent needs from its brain
A personal agent like Hark asks a lot of its model in one conversation.
It has to understand a messy request, ask the right follow-up question and then hand a clear plan to its computer.
My restaurant test showed that loop working well.
I typed, "Book an Italian restaurant for Wednesday next week."
Hark asked how many people and what time, and it understood my typo'd reply of "6pm, two people" without any fuss.
Then it started booking one of my favourite restaurants on its own computer screen.
That first half is the main model's job, and the second half is Handoff's job.
GoldieBench speaks to the first half, because the board rewards models that read a brief carefully and get the details right first time.
A model that averages 7.57 across 50 one-shot builds is very capable at following instructions, which is exactly what a booking request needs.
The board can't speak to the second half, which is why Hark's Handoff results and your own testing matter.
Hark's Handoff claim, read next to the board
Hark compared Handoff with Opus 4.8, so it's worth seeing where that model sits on GoldieBench.
Opus 4.8 averages 7.51 on the board across 47 tasks.Hark says Handoff beat it by 2 points on average in browser tests, which is a different kind of test from ours.
So the fair reading is this: Hark claims its own browser model beats a strong general model at browsing, while the general model is still a solid builder on our board.
Those two statements can both be true, because using a website and building one are different skills.
If you're choosing between agents for real web tasks, I'd give Hark's browser claims a fair trial on your own jobs rather than taking anyone's leaderboard as the final word.
Why the model question still matters to builders
For most people using Hark, the model doesn't matter, because they just want the booking made.
For builders it matters a lot, for three reasons.
First, the model decides how well Hark follows a long or fiddly instruction.
Second, the model decides how good the websites and documents it builds will be, and the reported Opus 5.5 label covers website building.
Third, an undisclosed model can change without notice, so the agent you test today may not be the agent you use next month.
That's why I'd treat Hark as a polished personal agent and keep the work where I choose the model myself.
If you want a model you fully control, the local models board shows which free models are good enough to run on your own machine.
Limits and safety for the Hark AI agent
Hark can act on your accounts, and my restaurant test showed it really will go and book.
I told it to stop before it finished, and that's the habit to build.
Hark's terms say you can give it standing permission for a category of transactions, so keep confirmations on until you trust it.
Hark says it never sells your data or shares it with advertisers, and its terms say content reviewed for safety can be used for training even after an opt-out.
It doesn't speak back, because voice is input only, and launch coverage said it was rolling out in the United States first.
Should you switch agents again?
A viewer listed Hermes, Muse, Dots, Grok, OpenBot and now Hark, and asked whether the differences are big enough to set everything up from scratch again.
Looking at the bench, the brains are close enough that I wouldn't switch for the model alone.
Switch only if Hark does a real job better than the agent you already run.
If you've never had a personal agent working, Hark is the easiest place to start.
Also On Our Network
- Agent Operatorsthe operator's step-by-step Hark setup
- AI Income Desknine business uses for the Hark AI agent
- AI Tool Verdictour feature-by-feature Hark review
- agentos.guidea Hark quick-start for your Agent OS
- aiprofitboardroom.comGrok Bot vs Hermes Agent on the AI Profit Boardroom blog
Until Hark names its model, judge it on real tasks and use the bench to understand the brains behind the Hark AI agent.
FAQ
What model powers the Hark AI agent?
Hark has not officially named the model behind its main assistant. TestingCatalog reported a Claude Opus 5.5 label in Hark's configuration for the main assistant, task agents and website building, but warned that labels do not confirm which model handled a request. Hark's browser agent, Handoff, is its own post-trained computer use model.
Is Hark on GoldieBench?
No. Hark is an app, not a model, and GoldieBench scores models on one-shot builds. Claude Opus 5.5, the model label reported for Hark's main assistant, scores 7.57 on GoldieBench across all 50 tasks.
Is the Hark AI agent free?
Yes. Hark says every feature is free. Hark Pro² is $20 a month for 2x more usage and Hark Pro³ is $100 a month for 10x more usage.
How does Hark compare with Grok Bot and Muse on the bench?
Grok 4.7 scores 7.15 and subscription Grok scores 8.09, Muse Spark 1.2 scores 7.55, and Claude Opus 5.5 scores 7.57. The brains are close, so setup and design matter more when choosing a personal agent.
Does GoldieBench test computer use agents like Handoff?
No. GoldieBench measures one-shot builds such as games, pages and simulations. It does not test browsing or bookings, so Hark's Handoff claims come from Hark's own Online-Mind2Web results.


