The future AI workforce will use subscriptions, APIs, and self hosted models at the same time
SGSeb Galindo, Founder
MachineHuman
cloud or local, a false choice.
The AI industry loves false choices.
Cloud or local. Subscription or API. Closed model or open model. Centralized or self hosted.
A real company will use all of them.
Different work has different requirements. Some jobs need the strongest model available. Some need low and predictable costs. Some depend on sensitive data. Some need to run thousands of times. Others require access to the files, tools, and permissions on a specific computer.
There will not be one correct way to power an agent.
The important question is not whether agents should run locally or in the cloud.
The important question is: Where should this work happen, and what intelligence should power it?
Three ways to power an agent
Companies already have access to AI through three different economic models.
Subscriptions
Teams already pay for products such as Claude, ChatGPT, Codex, Gemini, and Cursor.
Subscriptions make powerful models accessible at a predictable monthly cost. They are often the easiest way for a person or small team to begin using agents.
If a company already pays for that intelligence, its agents should be able to use it.
Direct APIs
APIs make sense when work needs to run at scale, inside a product, or through a highly customized workflow.
They let companies choose specific models, control how requests are made, and pay according to usage.
An API is not inherently better than a subscription. It simply has different economics and different operational advantages.
Self hosted models
Local and self hosted models give companies another option.
They can run on a laptop, a team server, a GPU, an AI box, or infrastructure inside the company’s own cloud account.
This can make sense when data is sensitive, usage is high, latency matters, or the company wants more direct control over the model and its costs.
Local does not have to mean one person running a small model on a laptop.
It can mean an entire organization operating its own intelligence infrastructure.
Figure 1. Three economic models, three agents, one team. Each agent uses what fits its work.
One company will use all three
Imagine a software company with several agents.
A coding agent uses an existing Claude or Codex subscription on a developer’s computer. It has access to the repository, terminal, development environment, and credentials already present on that machine.
A support agent runs through an API on an always on server because it needs to handle requests throughout the day.
A research agent uses a self hosted model on a company GPU because it works with sensitive internal documents.
A planning agent uses a different provider because that model performs better for long context reasoning.
These agents use different models, run on different computers, and have different economics.
They should still be able to work together.
The coding agent should be able to ask the research agent for context. The planning agent should be able to hand a task to the coding agent. A person should be able to follow the work, review the result, and step in when judgment is required.
The company should not have to rebuild its workflow every time a better model appears.
taskcontextreviewyou[o_o]codingclaude or codexsubscription[^_^]planninganother providerlong context[-_-]researchself hostedcompany gpu[o.o]supportapialways on server
Figure 2. One company, four agents. Different models, different computers, different economics, one flow of work.
Run the work near what it needs
Agents are most useful when they can reach the real environment where work happens.
That includes files, repositories, databases, browsers, applications, credentials, internal tools, and company permissions.
Sometimes the right environment is a cloud machine.
Sometimes it is a server in the office.
Sometimes it is the laptop where a person already has everything configured.
The execution environment should follow the work.
This is why the future of agents cannot depend on moving every file, credential, and workflow into one vendor’s cloud.
It also cannot depend entirely on employee laptops being open and connected.
Companies need the freedom to place each agent where it makes the most sense.
Every computer becomes potential workforce capacity
A company’s computers were designed for people to use directly.
Agents change that.
A laptop can support a personal coding or research agent. An office server can host shared agents for a department. A GPU can power local models. A cloud machine can keep scheduled and long running work active.
The machines a company already owns or controls become capacity for getting work done.
But distributed capacity creates a coordination problem.
Who can use each machine? Which agent runs there? What tools and credentials can it access? Who owns the agent? What happens when work needs to move to a different computer?
A collection of computers and agents is not enough.
The company needs a network that coordinates them.
[o_o][^_^][-_-]Laptop
[-_-][^_^][o.o]Server
[•_•][o_o][^_^]GPU
[^_^][-_-][o.o]Cloud
Figure 3. The machines a company already owns or controls become capacity for getting work done.
The orchestration layer should remain neutral
Cyborg is built around a simple principle: the company should decide how each agent is powered and where it runs.
An agent can use an existing AI subscription, a direct API, or a self hosted model.
It can run on a laptop, server, GPU, AI box, or cloud machine.
It can work with agents using completely different models and infrastructure.
The model can change without destroying the agent’s identity, job, memory, permissions, or place in the organization.
Cyborg does not need every agent to use the same provider.
It needs them to work as one team.
Choice becomes an organizational advantage
The best model today may not be the best model six months from now.
The cheapest way to run one workflow may be the most expensive way to run another.
A provider might change its pricing. A new open model might become good enough to run internally. A sensitive project might require different infrastructure from the rest of the company.
Companies that build their entire agent workforce around one provider will inherit every limitation and pricing decision that provider makes.
Companies that can move intelligently between subscriptions, APIs, and self hosted models will have more control.
They will be able to choose the right combination of quality, cost, privacy, speed, and ownership for each job.
The future is not cloud or local.
It is not subscription or API.
It is a coordinated workforce that can use all of them.
The intelligence can come from anywhere. The company should remain in control.