Researched October 4, 2026. Here’s how I would build an AI-powered solo company using tools available today. The capabilities are backed by sources; the operating model is a proposal, not a description of how Homann Software currently runs.

Here’s the morning I have in mind. It’s 7 a.m. I open my dashboard, coffee in hand. There’s a market brief with sources, a tested code change ready for review, a support queue with the tricky cases flagged, and a proposal draft that matches what I can actually deliver.

That’s the kind of morning I’d like to build a business around. I look after the customers, set the direction and make the final decisions. AI agents handle much of the repeatable work, including jobs that can move forward while I’m asleep.

In October 2026, we have useful tools to make this happen. Agents can research, work with files, investigate bugs and coordinate longer tasks. Getting useful work out of them is increasingly practical. Getting that work delivered reliably, without spending all day watching agent conversations, still takes careful design.

I’d start with one clearly defined service or software product. Each agent gets a specific assignment, clear limits on what it can do, and a way to show me that the job is done properly.

Why I’m paying attention to this now

For Homann Software, this gets interesting where AI meets Java, Spring and software architecture. Three recent developments are especially relevant:

Development Recent primary evidence What it changes for a founder
Agents doing substantial work under supervision Anthropic’s August 2026 internal measurement snapshot Delegation can cover large pieces of real engineering work; supervision remains part of the design.
Execution that survives interruptions Temporal’s September 21 AgentCore integration article Overnight work can retain progress instead of depending on one uninterrupted chat or process.
Better control over retrieved business knowledge Spring’s October 2 Modular RAG and TypeSafe Jev article Agent answers can use more selective evidence from company documentation.

The Temporal example addresses a very practical problem: a worker can disappear halfway through a job. Its design separates the agent runtime from the system that keeps track of progress. The Serverless Workers integration is still prerelease, so I’d evaluate it carefully before relying on it for customer work. Temporal integration announcement.

Spring’s example prepares better search queries, then filters and reranks the passages it finds. That could help an agent find the product or customer information it actually needs. There’s a detail I’d check before putting it into service: the Jev components let passages through if their service is unavailable. That behavior matters, especially if someone expects the filter to protect a business action. Spring engineering article, October 2.

For me, the solo company powered by AI agents is the most useful way to bring these developments together. It connects the technology to questions a founder cares about: What can I offer? What can I reliably deliver? And what does each job cost? That’s my editorial choice for Homann’s audience, rather than a claim about an industry-wide ranking.

Anthropic provides a striking example of how far delegation has come. The company reports that Claude led 26% of its AI R&D work in August 2026 and collaborated or did more on over 90%. None of the measured work operated fully autonomously. Its most-used internal platform had about 30,000 research and engineering agents active at a time, with online and offline monitoring. These are Anthropic’s own measurements in a specialized setting, so I wouldn’t use them to predict a small company’s productivity. Anthropic’s measurement report.

What I take from this is that substantial delegation is possible, and that the systems around the agents deserve just as much attention as the agents themselves.

What I would delegate first

I’d start with jobs where I can check the result and fix a mistake without causing a bigger problem. Here’s the shortlist I’d work from. These workflows combine available capabilities; getting them to work in a particular company still depends on its tools, connectors and account access.

Business function Useful work to delegate today Completion evidence Owner checkpoint
Market research Follow selected sources, compare competitors, flag relevant changes Dated sources, claim-to-source links, uncertainties Positioning and commercial decisions
Sales preparation Qualify inquiries against explicit criteria; draft discovery notes and proposals Customer requirements, assumptions, scope and estimate basis Price, commitments and sending the proposal
Software delivery Investigate an issue; implement a bounded change; run checks; prepare a reviewable diff Diff, executed checks, failure notes and reproduction steps Architecture tradeoffs and production release
Content Prepare a sourced article, diagrams, examples and distribution drafts Source register, tested examples, preview and image provenance Editorial voice, factual acceptance and publication
Customer support Categorize requests, retrieve approved answers, draft responses, identify escalation cases Relevant policy version, proposed answer and escalation reason Exceptions and unapproved commitments
Operations Inspect telemetry, reproduce alerts, prepare an incident brief and recovery proposal Observations, timeline, affected systems and proposed action Recovery beyond a tightly approved runbook
Back office Organize records, extract invoice fields and prepare reconciliation exceptions Source documents and discrepancies Payments, submissions and final accounting decisions

There are already tools for several of these jobs. OpenAI’s current Agents API documentation describes managed sessions, orchestration, recovery, sandbox execution and MCP connections. Its examples cover issue investigation, document review, read-only data analysis and incident investigation with approval requests. I’d still need to connect those capabilities to my own systems and processes. Official OpenAI documentation.

Anthropic’s Claude Tag announcement describes agents working through tasks with connected tools and scheduling follow-up work over hours or days. It launched in beta for Team and Enterprise customers in Slack. Longer-running delegation is becoming something you can buy and integrate, although I’d check access and deployment requirements before choosing a platform. Claude Tag announcement.

My first service might be a recurring architecture and maintenance review for a specific Java application. Agents collect information they’re allowed to access, identify changes, run an agreed set of checks and prepare a report. I review the findings and advise the customer. It’s a manageable place to start: a clear offer, a concrete deliverable and a result I can assess.

Give each agent a clear job

Calling an agent “Head of Sales” doesn’t tell it how to do the job. I need to define its inputs, the result I expect, the tools it can use and the point at which it should stop.

A useful assignment could read like this: review three competitor websites we’re allowed to access, cite the important findings, and save a draft brief in the company workspace. Stop after two attempts or when the assigned allowance runs out. Leave form submissions, prospect outreach and purchases to a separate, approved process.

I’d keep a small set of reusable roles: researcher, delivery worker, verifier, content worker and operations investigator. Each role can handle many jobs over time. There’s no need to keep a model running constantly for each one. A dispatcher starts jobs when an event or schedule calls for them and keeps track of the results.

Proposed architecture: owner decisions feed a durable dispatcher, bounded specialist workers, and restricted tool adapters; evidence returns to the owner.

Figure 1. A proposed operating model. The workflow owns state and transitions; workers produce artifacts within a defined scope.

The dispatcher needs a reliable record of where each job stands. When agents hand work to each other, I’d have them link to the original, versioned files. Repeated summaries can lose details. The customer’s agreed scope, a tested commit and the current support policy should each have a clear home and someone responsible for them.

This fits how I’d approach engineering at Homann: clear domain boundaries, useful use-case models, deliberate UX/UI choices, explicit quality requirements, tests and ADRs. I’d keep improving them as working software and customer feedback reveal what’s missing. Agents can help connect a requirement to the code change and the checks behind it. I still need to decide whether the result is good enough to accept.

Interoperability helps connect the pieces. MCP standardizes connections between AI applications and tools or data. A2A addresses communication between independent agents, including discovery, delegation and results. I would use a framework’s ordinary internal calls for workers in one application, and consider A2A when independent agent services genuinely need to cooperate. MCP introduction, A2A documentation.

Once those connections work, my application still has to check whether a requested action is allowed for this customer, this budget and this job.

What “24/7” should mean in practice

For me, round-the-clock work means the system is available whenever there’s a useful job to do. It works through a defined queue, starts when a schedule, customer event or alert calls for it, and waits when it needs more information or approval. When the allowance runs out, it stops.

A loop that keeps inventing new tasks can burn through money quickly. I’d give the overnight shift a simple routine:

  1. At 18:00, select jobs, record approved scope and inputs, and allocate the shift’s budget.
  2. During the night, allow authorized research, draft creation, isolated code changes and test runs. Retry only failures known to be safe to retry.
  3. On a sensitive proposed action, record the exact payload and artifact revision for morning review. Continue unrelated authorized jobs.
  4. At 07:00, show completed artifacts, failed jobs, budget consumption and decisions awaiting the owner.

Overnight operating cycle: evening contracts, bounded overnight work, paused consequential actions and morning evidence review.

Figure 2. An illustrative shift, not a promise about elapsed task time. Useful work continues while decisions requiring the owner remain pending.

In the morning, I want a short briefing: what changed, which checks passed, what’s still uncertain and what needs my decision. I don’t want to spend breakfast reading a hundred transcripts.

The system also needs to remember where it left off. Temporal documents how workflows recover their state and replay recorded events. In production, I’d store jobs and pending decisions so work can resume after an outage. Model calls and external interactions need to run through the runtime’s appropriate activity boundary. Temporal execution documentation.

There’s a catch when a job changes something outside the workflow. A provider might accept a request just before the response gets lost. Trying again mustn’t create a second effect. Homann’s tested Java article on durable workflows and idempotency looks at that problem in detail.

The practical details matter too: the host has to be available, credentials have to work, tools have to be reachable, and rate limits need handling. If the desktop is asleep, a saved schedule may be of little use. Before relying on unattended work, I’d check how the chosen product actually hosts and starts jobs, and make sure the infrastructure is monitored.

A small Java example: deciding which jobs can run

Here’s a small Java 21 example that makes one part of this design concrete. It checks whether a job can start and reserves its share of the allowance before running it.

The contract contains an action, an admission deadline, at most three attempts and a fixed number of abstract budget units per attempt. Before starting, the dispatcher reserves enough units for every permitted attempt. Started attempts remain charged at that conservative allowance; unstarted attempts are released.

The units here are a simple way to control which jobs can start. They don’t measure API spend. A production adapter would also have to enforce model and tool limits for each attempt, then reconcile actual provider usage. Without that, a job could pass this check and still cost more than expected.

The core checks are ordinary application code, outside the model:

        if (stopped || !scope.contains(job.action()))
            return new Result(Status.DENIED, 0, "Stopped or outside worker scope");
        if (!clock.instant().isBefore(job.deadline()))
            return new Result(Status.EXPIRED, 0, "Admission deadline reached");
        if (REVIEW.contains(job.action()))
            return new Result(Status.REVIEW, 0, "Owner must review the exact proposed action");
        if (!budget.reserve(job.reservation()))
            return new Result(Status.BUDGET, 0, "Shared shift budget exhausted");

In this example, sensitive actions return REVIEW without calling the tool. There’s no shortcut for approval. A production approval service would check the owner’s identity, record consent for the exact action and file revision, check that it’s still valid, and pass the command to an executor authorized to carry it out.

Several workers may ask for the remaining allowance at the same time. This check handles that inside one process:

        public synchronized boolean reserve(long units) {
            if (units <= 0) throw new IllegalArgumentException("Nonpositive reservation");
            if (units > limit - allocated) return false;
            allocated += units;
            return true;
        }

Using subtraction avoids an overflow before we compare the numbers. Synchronization makes sure threads sharing this Budget instance can’t reserve the same remaining allowance twice. Workers in separate processes would need a shared transactional store.

Run the source and tests with a JDK 21 or later:

javac --release 21 -Xlint:all -d out NightShift.java NightShiftTest.java
java -cp out NightShiftTest
java -cp out NightShift

The executed demo produces:

research=DONE
draft=DONE
publish=REVIEW
allocated_units=25

The demo uses fixed tool responses so you can reproduce the result without a model account. It doesn’t do real market research. The order of work is enforced by the application: if research fails, drafting doesn’t start; if the draft fails, publication isn’t proposed.

Verification: 24 tests passed on Temurin OpenJDK 21.0.11+10. The suite covers sensitive actions before tool invocation, worker scope, exact deadline boundaries, unused reservations, retry exhaustion, permanent failures, interruptions, invalid contracts and overflow. In the concurrency test, 20 workers compete for 50 units: five 10-unit attempts run and fifteen are rejected for budget exhaustion.

The example does not test a model, a cloud scheduler, connectors, persistence, cross-process coordination or actual billing. Deadlines and stop requests prevent admission and subsequent attempts; they cannot interrupt a tool call already running. Adapters need enforceable timeouts and cancellation. Retrying a tool is appropriate only when its operation is safe or protected by receiver-enforced idempotency.

Download the runnable example and verification records. The source package includes reproduction instructions and all test output.

How much authority should an agent have?

I’d use three operating modes in this company. They’re choices about how I want the business to run; individual products may have different rules.

Prepare independently: read approved sources, create drafts, run tests in an isolated environment and inspect authorized telemetry. The agent can get useful work ready without making commitments on my behalf.

Act within an agreed policy: carry out a specific action with an approved target, limits, conditions and failure plan. For example, an agent could restart one staging service under an agreed runbook. The adapter checks that the actual request still fits what was authorized.

Ask the owner: bring me a new commercial commitment, publication, payment, sensitive customer exception or production change outside the runbook. Show me exactly what would happen so I can make the decision.

Authority ladder: prepare unattended, execute within a previously approved narrow policy, or wait for an owner decision on a consequential action.

Figure 3. Increase autonomy per action after evaluating evidence and failure consequences. A role title grants no authority.

Checks belong beside the tools that create effects. OpenAI’s documentation explicitly warns that agent input and output guardrails do not cover every custom tool call in a manager workflow. Guardrails and human review.

There are also new infrastructure controls in this area. AWS announced stateful temporal policies and traffic rate limiting for Bedrock AgentCore in August 2026. Those temporal policies use prior interaction events when evaluating access; they are distinct from Temporal’s workflow engine. They illustrate the move toward enforcing policy in infrastructure as well as in prompts. AWS announcement.

A second agent can help review a proposal. If both models agree, I’d still want to see what they checked. Give the verifier the original evidence and deterministic checks wherever possible. I’m responsible for the company’s promises, so each delegated job needs a clear definition of a satisfactory result.

Keep an eye on the economics

The business appeal is easy to see. A small company may be able to offer more without taking on a large payroll. Whether that pays off depends on customer demand, delivery quality and how much time I spend reviewing the work.

Anthropic’s research-system account reports roughly fifteen times the token use of ordinary chats for its multi-agent setup and warns that some tightly coupled work is a poor match for parallel agents. That is a workload-specific observation, not a universal cost multiplier. I would begin with one worker and add parallelism where it demonstrably improves accepted outcomes. Anthropic’s multi-agent engineering report.

I’d watch five things: what an accepted deliverable costs, how many minutes I spend reviewing it, how often it needs rework, how long customer jobs have been waiting, and whether we’ve had serious incidents or policy violations. Token usage helps explain the bill. It doesn’t tell me whether the customer got something useful.

Let’s put some hypothetical numbers on it. An overnight run costs EUR 12. I spend 25 minutes reviewing it, valuing my time internally at EUR 80 an hour, and allow EUR 15 for expected rework. That comes to about EUR 60.33 before infrastructure overhead. If doing the same job manually takes two hours at that rate, the difference is EUR 99.67. These are illustrative assumptions, not measured savings or a pricing recommendation. The calculation only helps if the work meets the acceptance criteria and a customer actually wants it.

I’d be equally careful with productivity headlines. METR’s February 2026 update explains why selection effects and developers using agents concurrently made its newer experiment unreliable for estimating current productivity gains. Older slowdowns and newer self-reported gains both need context. I’d measure the effect on my own work before promising a multiplier. METR’s experiment update.

How I’d spend the first month

Week one: pick one deliverable customers will pay for and write down how I produce it today. Record review time, common problems and what a good result looks like. Turn that into a clear assignment and gather a few representative cases to check the agent against.

Week two: let one agent prepare drafts alongside the existing process, without acting on them automatically. Compare the results. Fix missing context, unclear use cases and weak tests before handing it more responsibility.

Week three: connect a durable job queue, restricted adapters and a shared budget. Start with preparation work overnight. Check what happens during outages, when credentials expire, when a job arrives twice, when inputs change and when I’m unavailable.

Week four: look at the results. If the checks support it, allow a few specific actions under an agreed policy. Keep important exceptions visible. Add more work when the quality and review effort make it worthwhile.

I’d also make sure I can pause the system when I’m unavailable and recover without having to remember which agent was doing what. In a solo company, a failed handoff comes back to me. A good operating setup keeps that manageable.

The company I’d want to run

I like the idea of staying small while being able to get a lot done. I’d build that capacity around a clear offer, useful company knowledge, work I can check, and a dependable path from the customer’s request to a result they can use.

Agents can already help with much of that work. My job as founder would be to put the pieces together so their output becomes dependable enough to sell. Each morning, I’d want fewer unresolved tasks and a clearer view of what needs my attention.

That’s a solo company I’d be happy to run: plenty of work moving forward, clear limits on what agents can do, and customers who can rely on what we deliver.

Further reading: Homann’s AI agent harness architecture and durable workflow receiver example. This article’s code establishes only the admission-policy properties described above; no end-to-end autonomous business was deployed or evaluated.