Hopp til hovedinnhold
AIKI

AI agents complete 16 percent of freelance jobs at pro level

||5 min lesing

Key takeaways

  • Remote Labor Index measures 240 real freelance projects worth 144,000 US dollars in total value
  • Claude Fable 5 leads at 16.1 percent, twice as high as the next best model
  • In 8 months the share of projects an AI agent can handle at pro level has gone from 2.5 to 16.1 percent
  • Source criticism: AI judges are 2.5 to 3 times too generous, human evaluation is required
  • For Norwegian SMBs this means around 1 in 6 typical office projects can be automated today

On July 1, 2026, the Center for AI Safety released a new version of the Remote Labor Index, the benchmark that measures how often AI agents can complete real, commercially valuable freelance projects at a quality level a paying client would actually accept. The answer: 16.1 percent. Eight months ago the number was 2.5 percent. That is a sixfold increase.

What is the Remote Labor Index

The Remote Labor Index is developed by Scale Labs in partnership with the Center for AI Safety. The test uses 240 projects sourced from 358 verified freelancers, with a combined value of 144,000 US dollars. The projects cover seven categories: 3D and CAD, architecture, graphic design, video and animation, audio, data analysis, and web applications.

Each project is given to an AI agent that has access to a standardized Linux machine with 30 professional tools, including Blender, GIMP, and Audacity. The agent can use up to 24 hours of compute per project. A second AI agent acts as a critic in a loop: the second agent reviews the output critically, and the first agent revises the work. Finally, human evaluators from CAIS score the deliverable against a gold standard created by a professional freelancer.

The measured metric is the «automation rate», the share of projects where the AI deliverable is judged at least as good as the human gold standard.

Who leads the pack

Claude Fable 5 tops the list with 16.1 percent. That is twice as high as Opus 4.8 at 8.3 percent, which in turn is twice as high as its predecessor Opus 4.6 at 4.17 percent. GPT-5.5 scores 6.3 percent. Gemini 3 Pro, a newer model, sits near the bottom at 1.25 percent.

Important context: 22 of 240 projects were not evaluated for Fable 5 because US authorities have export restrictions on the model. Even in a worst-case scenario where the agent failed on all of these unevaluated projects, Fable 5 would score 14.6 percent, still above every other model.

Progress does not follow release date. Gemini 3 Pro is one of the newest models in the test and still scores near the bottom. Fable 5 is the oldest model that scores highest. That means it does not help to buy the newest model. It is model-specific agent performance that matters, and the differences are large.

Where 16 percent actually hits

It is tempting to read 16 percent as «AI can do every sixth task». That reading is wrong. RLI measures a narrow segment: projects where an AI deliverable is good enough that a customer would pay for it. That is a much higher bar than «AI can produce something usable».

At the same time it is a lower bar than «AI is as good as a human». The benchmark says nothing about price, speed, or the ability to build long-term customer relationships. An AI agent uses 24 hours of compute per project, and for many assignments that is significantly more expensive than hiring a freelancer.

So what actually hits the 16 percent threshold? Typically: first drafts of designs, data visualizations, prototype web apps, short marketing videos, audio editing, and formatting. What does not hit: assignments that require cross-disciplinary understanding, customer dialogue, sales, physical presence, or iterative creative conversation with the client.

An important source criticism

CAIS tested whether AI judges can replace human evaluators. The answer is a clear no. When AI judges scored the same deliverables as humans, they were shockingly credulous. For GPT-5.5 the AI judge scored almost three times too high. For Opus 4.8 it was two and a half times. The ranking came out right, but the numbers were way off.

The reason is that proper evaluation requires opening files in professional software and actually using them the way a paying customer would. That is exactly what AI agents are worst at, and it is the same thing that limits AI judges.

The examples from the CAIS report show why this matters. On a «ring design» task Fable 5's result looked professional at first glance, but fell apart on closer inspection. On an architecture project GPT-5.5 faked an appealing 3D render while the actual 3D model remained flawed. These are the kinds of errors humans catch, and which AI judges miss.

What this means for Norwegian SMBs

For Norwegian businesses with 7 to 100 employees the relevant question is not «does AI take 16 percent of jobs», but «which 16 percent of my repetitive tasks can an autonomous agent do at good enough quality that I don't have to do them manually».

A concrete example. A small accounting firm receives around 50 receipt PDFs per week. Today the accountant manually types in vendor, date, amount, and VAT category. With an autonomous AI agent, for instance one of the solutions we build in Autonomous AI Agents, the agent can run OCR, extract the fields, categorize against the chart of accounts, and send a booking proposal. The human opens the proposal, corrects exceptions, and approves.

The entire shift from 2.5 percent to 16 percent automation rate means in practice that an AI agent today can take the first draft of these kinds of tasks. It is not a replacement for the accountant. It is a tool that frees up 12 to 15 hours per week for a seven-person firm, enough to take on more customers without hiring more staff.

For other industries the picture looks similar. A car dealership can let an agent handle the first draft of offer emails while the salespeople focus on customer meetings. A logistics company can automate report generation while the controller focuses on analysis. A hospitality business can let agents answer standard emails while the hosts concentrate on the guests.

If you want to map out which processes in your business are ripe for this kind of automation, an AI Audit engagement is a natural starting point. We go through the workflow over 6 weeks and deliver an ROI matrix where you see exactly which projects hit the 16 percent threshold today, and which require more mature technology.

What does not hit yet

Two things are worth keeping in mind.

First: 84 percent of projects in RLI are still outside AI reach. That is a conservative measure. When you add real Norwegian work tasks that require customer dialogue, government contact, or physical presence, the share of AI-suitable work goes down. For most Norwegian SMBs that means 70 to 80 percent of work hours still require humans.

Second: AI agents are not free. They require setup, maintenance, and human quality assurance. An AI Partner agreement covers operation and development over time, but it is not something you set up once and forget. AI Automation of a single workflow can deliver immediate value, but it requires someone to own the solution over time.

A sober perspective

The CAIS benchmarks are honest. They use human evaluators, they state the worst case for Fable 5, and they explicitly warn against AI judges. That makes the numbers more credible than most «AI is 10x better» claims you read on social media.

At the same time it is worth remembering that the Remote Labor Index only measures one aspect of AI capability: the ability to deliver a finished, professional deliverable on a one-off assignment. It says nothing about how AI agents perform as collaborators over time, how they handle unexpected problems, or how they build trust with customers.

For Norwegian SMBs the conclusion is still clear: the time window for adopting AI agents on simple, repetitive tasks is open now. Those who wait will find that the tasks they thought were impossible to automate can actually be automated tonight. Those who start build the competence needed when the automation rate passes 30 percent, and that day is coming faster than most expect.

Want to know which processes in your business are ready for AI agents today? Book a free mapping session, and we will go through three concrete workflows and see what can be automated.

Del:LinkedInXFacebook