Researched October 7, 2026. A business and architecture perspective for founders who use agents across delivery, sales, support and administration. The numerical scenarios are illustrative; they are not Homann Software operating results or provider quotations.
The demo ends when the agent produces a pull request. Your business day does not.
Someone still has to qualify the lead, understand the customer, price the work, check the proposal, maintain the software, answer support questions and prepare the paperwork. For a solo business, that someone is usually the owner. AI agents can help with all of this. The interesting question is whether their work releases usable capacity at an acceptable cost.
That is the missing conversation behind many “one person, a whole company of agents” stories. An impressive output is easy to show. The invoice, rejected drafts, review queue and months of maintenance are harder to fit into a social media clip.
Our earlier article, The Solo Founder’s AI Company, explored how such work could be orchestrated. This article asks the commercial question: what would make that operating model worth running?
Here, token economics means the economics of AI usage. It does not mean issuing a cryptocurrency token, and calling an agent an “employee” does not make it a substitute for every responsibility an owner carries.
Why this is the timely topic
Homann Software’s focus on software architecture and Java engineering makes reliable AI operations a natural topic. The freshest relevant signals point toward costs and accountable outcomes:
- On October 2, 2026, Gartner published a FinOps playbook for custom-built agents. Its public abstract connects agent economics with total cost, value and runtime control. The full report is paywalled; this article draws only on that abstract.
- Google’s August 26 announcement introduced billing flexibility and cost controls for agent workloads. This is a vendor announcement, not evidence that any particular solo business becomes profitable.
- The State of FinOps 2026 puts AI prominently on the technology-value agenda. Its enterprise-oriented findings establish context, not a representative survey of sole proprietors.
- In the Java ecosystem, Spring’s October 2 Modular RAG example focuses on retrieving material and retaining what answers the question. Selective context is relevant to efficient engineering, although that article does not establish savings for the business scenarios below.
My editorial choice is agent unit economics: it connects those developments with decisions a small business owner can make today. This is a relevance-based selection, not a claim that one topic objectively outranks the entire IT market.
Why build an agent-assisted business in the first place?
For a solo founder, the attraction is straightforward: more useful work from a limited budget and a limited number of personal working hours. The owner can spend more time understanding customers, shaping the offer and delivering value while agents prepare research, assemble proposal drafts, investigate support cases and handle repeatable administrative steps.
Better cost-benefit efficiency is the first potential advantage. A well-bounded workflow can reduce the effort needed to produce an accepted result. The commercial opportunity is especially attractive when several business functions need occasional capacity but do not each justify a dedicated role. Whether that advantage survives review, integration and upkeep is what the calculations in this article are designed to reveal.
More productive, value-creating output is the second. A researched proposal ready for a customer conversation, a resolved support case or a tested improvement can move the business forward. Agents can prepare independent tasks in parallel and continue authorized work while the owner is occupied elsewhere. The gain comes from completed work that customers or the business can use; generating more drafts alone does not establish it.
Less internal friction from personal agendas is a third attraction, and one that receives less attention. An owner may want working capacity without also creating another layer of status competition, territorial behavior or interpersonal negotiation. An agent has no personal career to defend, promotion to pursue or resentment about another worker receiving credit. Reassigning a bounded task, requesting another attempt or discarding a weak draft need not become a discussion about rank or wounded pride. That can make the working process more direct and leave more attention for the customer.
This is an operating-model argument, not a measured guarantee of higher productivity. AI can still produce contradictory answers, defend a mistaken conclusion in its wording or agree too readily with the owner. Poor instructions can create waste just as poor incentives can. Independent verification and explicit acceptance criteria remain necessary. Human colleagues also contribute trust, initiative, tacit knowledge and constructive disagreement; a founder should preserve access to those strengths through customers, peers and specialist advisers.
The ambition is a business with more usable capacity and fewer avoidable coordination burdens. It still needs a clear offer, customer relationships and an owner willing to make decisions. The economics become compelling when agents help that owner produce better accepted outcomes with less total effort—and when the released capacity is put to worthwhile use.
Start with an accepted result
“Cost per million tokens” is useful when comparing inference rates. It is insufficient when comparing workflows.
A proposal agent might use an inexpensive model but produce drafts that need twenty minutes of correction. A more expensive workflow might produce fewer mistakes and require five minutes of review. Neither comparison is complete until the same acceptance standard has been applied.
Define the unit first:
| Workflow | Count as accepted when… | Owner responsibility that remains |
|---|---|---|
| Research | The brief answers the question and its key claims have checked sources | Decide whether the evidence supports a business decision |
| Sales proposal | Scope, price assumptions and exclusions have been reviewed | Approve commitments and send the proposal |
| Support | The resolution is correct and the case needs no further work under the agreed criteria | Handle sensitive cases and customer accountability |
| Receipt preparation | Original documents are linked and uncertain fields are flagged for review | Review classification and the accounting handoff |
| Coding | The change satisfies the use case and required checks | Accept delivery and operational consequences |
These are proposed acceptance rules, not claims about mandatory legal procedures. A drafted answer is not automatically a resolved case; a prepared receipt is not a completed filing. Use the actual endpoint of your workflow.
For a batch of work, I would track two figures:
Cash per accepted result = all attributable cash expenditure / accepted results.
Economic cost per accepted result = (cash expenditure + valued owner time) / accepted results.
Include the expenditure and review time of failed attempts in the numerator. Otherwise, the workflow looks better precisely when it wastes more effort. If no results are accepted, report the cost and zero accepted results; the unit cost is undefined.
The FinOps Foundation’s AI overview provides useful context for cost allocation, forecasting and value measurement. The formulas here are a deliberately small operating model for a solo business.
Tokens are one layer of the bill
An agent’s model use can include system instructions, tool schemas, retrieved documents, growing conversation history and several attempts at the same task. Multiple workers and independent verification add more calls. A failed call can still consume chargeable resources.
Use the provider’s actual billing categories. Uncached input, cache reads, cache writes and output can have different rates. Do not count the same input twice, and do not assume a cache hit simply because a document was sent before.
Anthropic’s current pricing documentation distinguishes caching operations, asynchronous batch pricing and tool-related charges. Discounts and modifiers depend on the model, feature and deployment path. This is why the example below accepts a rate card rather than embedding a supposedly universal current price.
Then add the surrounding expenses: search and data access, execution environments, storage, monitoring, integrations and subscriptions. Allocate shared monthly costs consistently. Buying three platforms that overlap is still three bills; allocating a subscription should not make it disappear from your total.
Finally, count setup and maintenance: changing connectors, broken permissions, evaluation examples, prompt revisions and recovery after failures. During a pilot, record initial setup separately. When comparing ongoing operations, either amortize that investment over a stated useful life or show it alongside recurring costs. Do not hide it.
A small monthly scenario, including the owner's time
Consider five workflows in a hypothetical solo service business. The calculator uses an illustrative rate card of 3 currency units per million uncached input tokens, 0.30 per million cached input tokens and 15 per million output tokens. Owner time is valued at 90 units per hour. All cash inputs use the same currency; the example makes no foreign-exchange or tax assumptions.
The complete CSV includes aggregate token usage, paid tools, allocated recurring costs, review and upkeep minutes. Usage totals include retries and unsuccessful jobs. Initial setup, broader business overhead and remaining manual handling of unaccepted cases are outside this comparison and must be budgeted separately.
| Monthly workflow | Attempted / accepted | Model cash | All allocated cash | Economic cost per accepted result | Owner hours released |
|---|---|---|---|---|---|
| Research briefs | 40 / 32 | 5.16 | 51.16 | 13.41 | 9.13 |
| Proposal drafts | 20 / 12 | 6.60 | 51.60 | 36.80 | 4.67 |
| Support cases | 100 / 85 | 3.90 | 40.90 | 6.04 | 6.08 |
| Receipt preparation | 80 / 72 | 2.22 | 30.22 | 4.36 | 1.65 |
| Coding changes | 12 / 8 | 21.30 | 81.30 | 77.66 | 6.00 |
Rounded display values from the included, executed calculator. These are scenario results, not observed agent performance.
The proposal workflow illustrates the trap. Model consumption costs only 6.60, but tools and allocated fixed costs bring cash spending to 51.60. Review and upkeep consume 260 owner minutes, valued at 390. The total is 441.60 for twelve accepted drafts: 36.80 each.
Preparing those same twelve drafts manually at an assumed 45 minutes each values owner time at 810. The difference is 368.40. That is a modeled economic advantage under these assumptions, not cash profit or proven productivity.
Sensitivity matters more than the attractive first result. If proposal review and upkeep rise from 260 to 540 minutes, they consume all the time allocated to the manual baseline. The AI workflow then costs 861.60 economically, including its 51.60 of cash expenditure, compared with 810 manually. It has released no owner time.
Or keep the expenditure and review time unchanged but accept only six drafts. Economic cost per accepted draft doubles to 73.60. Investigate the acceptance rate before celebrating cheaper tokens.
Saved time becomes revenue only through another step
A solo business has a different constraint from a large employer: reducing the owner's administration time rarely removes a payroll expense. It may create time for paid delivery, sales, rest or more reliable service. Each can be valuable, but they are different outcomes.
Separate three measures:
- Cash expenditure: money that leaves the business for AI and its supporting services.
- Capacity released: measured manual effort avoided, less review, recovery and upkeep.
- Realized commercial value: additional work actually sold and delivered, contribution earned, or expenditure actually avoided.
Do not book the same released hour once as a saving at 90 per hour and again as new revenue. For a cash decision, compare incremental contribution or actual avoided payments with incremental cash expenditure. For an economic decision, compare complete alternatives consistently, including owner effort.
The illustrative workflows release about 27.53 owner hours in total. If only eight become additional paid hours at an assumed incremental contribution of 100 each after other incremental delivery costs, the result is 800 before the 255.18 of modeled AI cash spending. The remaining 544.82 is a scenario contribution before omitted setup, overhead and taxes. If none of the released hours produces additional business, there is no new revenue to count.
This is why a founder should ask “what will I do with the released capacity?” before buying another agent platform.
The surrounding business work needs an operating contract
For each workflow, write a small contract: permitted inputs, expected artifact, acceptance criteria, authority, resource allowance and escalation conditions. Budget time for keeping it current.
For research, that means dated evidence and checked originals. For proposals, it means an approved offer catalog, clear exclusions and a review of commitments. For support, it means access boundaries and a clear handoff when customer history is incomplete. For receipt preparation, preserve originals and route uncertain fields to the responsible reviewer. For publishing, it means source checks, editorial review and an explicit publication decision.
These boundaries also affect economics. Excessive approvals create a queue; broad uncontrolled authority creates expensive mistakes. Prefer unattended preparation where mistakes are recoverable, then put review close to consequential actions.
An operational spend control needs more than a dashboard. Before starting a task, reserve a conservative allowance for its bounded calls and tools. Parallel workers must share the same ledger; ten workers each observing “20 remaining” can overspend it together. Reconcile actual usage afterward, and hold unresolved usage until reconciliation succeeds.
Unknown cost should trigger a hold or an explicit exception. A timeout should not automatically release the reservation if the provider may still be working. Limit retries, tool fan-out, execution duration and output size. Reconcile local estimates with provider invoices because reporting delays and billing categories can differ.
This is a proposed architecture, not a tested production payment gateway. A local calculator cannot enforce a provider spending cap. Provider-side limits and restricted credentials are additional controls, and their scope must be checked for the chosen service.
A runnable calculator instead of a fictional ROI claim
The accompanying Python example uses Decimal for money arithmetic and the standard library only. Python keeps this business calculation easy to reproduce without a framework; the same model can sit behind a Java or Spring workflow ledger.
The token function makes the billing categories explicit:
def token_cost(uncached, cached, output, rates):
# Categories are disjoint; cache writes belong in separately rated input.
return (D(count(uncached)) * rates.input
+ D(count(cached)) * rates.cached
+ D(count(output)) * rates.output) / D("1000000")
The workflow function retains costs from unsuccessful work, adds owner review and upkeep, and divides only by accepted results. It also calculates a manual owner-time baseline for the same number of accepted results. In this demonstration, that baseline has no separate manual tools or cash costs; add them when they exist.
def workflow(row, rates, owner_hourly):
attempted, accepted = count(row["attempted"]), count(row["accepted"])
if attempted == 0 or accepted > attempted:
raise ValueError("Expected attempted > 0 and accepted <= attempted")
# Token totals, tools and review include unsuccessful jobs and retries.
model = token_cost(row["uncached_tokens"], row["cached_tokens"],
row["output_tokens"], rates)
cash = model + nonnegative(row["tools_cash"]) + nonnegative(row["fixed_allocated"])
review = nonnegative(row["review_minutes"])
upkeep = nonnegative(row["upkeep_minutes"])
owner_cost = (review + upkeep) * nonnegative(owner_hourly) / D("60")
manual_minutes = nonnegative(row["manual_minutes_per_accepted"])
baseline = D(accepted) * manual_minutes * nonnegative(owner_hourly) / D("60")
total = cash + owner_cost
return {"workflow": row["workflow"], "attempted": attempted, "accepted": accepted,
"model_cash": model, "cash": cash, "owner_cost": owner_cost,
"economic_total": total,
"cash_per_accepted": cash / accepted if accepted else None,
"economic_per_accepted": total / accepted if accepted else None,
"manual_owner_baseline": baseline,
"economic_difference": baseline - total,
"owner_hours_released": (D(accepted) * manual_minutes - review - upkeep) / D("60")}
From the extracted examples directory, run:
python -m unittest -v test_economics.py
python economics.py
The executed suite passed 17 tests on Python 3.13. They check mixed token categories, exact costs, failed-job allocation, owner upkeep, zero accepted results, invalid inputs, duplicate workflow allocations and sensitivity to rejection. The demo reproduces the table above. Download the source and reproduction instructions.
These checks establish the calculator's behavior for those cases. They do not validate the workload assumptions, live model quality, provider integration or a real business return. In production, normalize provider-specific usage into disjoint categories, version the rate card, include cache-write rates where applicable, and record request IDs for reconciliation. Avoid copying customer content into cost logs when identifiers and usage totals suffice.
What I would measure during a first month
I would start with two narrow workflows: one internal, such as a source-backed market brief, and one close to revenue, such as proposal preparation. I would measure manual work first using the same acceptance criteria, then compare the assisted process on representative cases.
For each job, keep the workflow and artifact revision, attempt count, usage categories, tool costs, review minutes, acceptance outcome and any later rework. Track the median and high-cost tail separately. Averages can hide a small number of jobs that consume the overnight allowance.
Review the results weekly. Did the output survive real customer use? Did exceptions create more interruption than the original process? Did subscriptions match the workload? Did released capacity become useful work? Expand only when both the artifacts and the economics hold up.
A small firm does not need a grand “AI organization” on day one. It needs a repeatable path from a bounded task to an accepted result, with costs and responsibilities visible. That is where architecture and business judgment meet—and where an impressive agent demo can begin to become a dependable company capability.