# The same task, 18 times the cost

*The hours were real, the output was correct, and the process still had to change.*

**TLDR:** A document-formatting task took four and a half hours. With the same brief and an AI agent, it took fifteen minutes for comparable output. The comparison changed two Holdex rules.

**Authors:** [Vadim Zolotokrylin](/people/vadim-zolotokrylin), Angelica Willianto

---

A routine task came back with four and a half hours against it.
The work involved moving a document between formats,
then cleaning up everything the formatter complained about.
The result was correct, the hours had been recorded honestly,
and there was nothing on the timesheet that looked unusual.
That was why the number deserved attention:
bad work would have given the reviewer something obvious to reject.

We repeated the same brief ourselves,
giving the mechanical part to an AI agent and checking the result.
Comparable output took fifteen minutes, one-eighteenth of the time.
Nobody had padded the time,
and the person who completed the original task had done what we asked.
The gap came from the method.

## The expensive method passed review

A normal review can check whether the work is correct and
whether the hours are plausible.
This task passed both checks.
Without the comparison, we would probably have approved it
and carried the same method into the next task.

The fifteen-minute version did not rely on unusual domain expertise.
The agent handled the document conversion and the formatter’s complaints,
while a person reviewed the result and decided whether it was ready.
A person still judged the output;
the repetitive execution no longer consumed their time.

That comparison changed the question.
“Were the hours real?”
was no longer enough.
We also needed to ask whether the method was reasonable
for the result it produced.
When acceptable work can consume 18 times more time because of how it was done,
effort alone cannot justify the cost.

## A company-level shift made of ordinary tasks

Companies routinely examine visible costs such as salaries,
software subscriptions, infrastructure, and partner rates.
Working method does not arrive as a line item;
its cost is distributed across tasks that look reasonable on their own.
One four-hour task will not change the company,
but the same difference repeated across mechanical work will.

a16z makes the broader company-level argument in
[There are only two paths left for software](https://a16z.com/there-are-only-two-paths-left-for-software/):
software companies need either to create enough growth through AI-native
products or to rebuild for much higher operating margins.
The paths differ, but both require AI to change the shape
and cost structure of the company.

a16z is talking about whole software companies;
we found a smaller version of that shift inside a document-formatting task.
Company-level cost curves are made of ordinary decisions like this one.
The gain appears only when the process changes,
not when someone works harder inside the old process.

## Judgment stayed with the person

Some work should take time because judgment, context,
and careful thought are the work.
We have written before about why AI is
[an excellent helper and a poor leader](/insights/ai-amazing-helper-terrible-leader).

In this task, the agent did not decide what a good document looked like.
It handled the mechanical steps, while a person defined the task,
checked the output, and decided whether it was ready.
The faster method preserved the judgment and removed the repetition.

## What changed at Holdex

The comparison led to two changes in how we operate.
First, we now evaluate work by the value it delivers relative to what it cost to
produce, rather than treating effort as evidence of value.
A day spent designing a valuable solution can be time well spent;
several hours spent on a mechanical process require a better explanation.

The second rule is more demanding.
Access to a paid AI agent, and the ability to get useful results from it,
is now a baseline expectation for our work.
We treat the subscription as basic working equipment,
like an internet connection or a desk, rather than a reimbursable add-on.

Credit limits are part of the cost question too.
When a plan runs out, we first ask
whether more capacity would unlock valuable work or fund an inefficient method.
We look at how requests are structured, how context is managed,
and what gets delegated before treating more credits as the answer.

We wrote both expectations into our operating rules so
[company decisions survive beyond the conversation that produced them](/insights/flat-companies-run-on-rules).

## The same timesheet means something different now

The useful comparison is selective.
For a task that is mostly mechanical,
one comparison can tell us whether the old cost still makes sense.
That comparison includes the time needed to prepare the instructions,
review the output, and correct what the agent gets wrong.

The original four-and-a-half-hour timesheet was honest,
and the work was correct.
We would still question the method now.
Once comparable output takes fifteen minutes with the judgment intact,
four and a half hours is no longer a neutral record of effort.
It is evidence that the process has not caught up.
