Meta says its personal agent Muse has a computer of its own in the cloud. It can browse websites, fill in forms and ask for approval before making a purchase. That raises an interesting question: what changes when we give an AI somewhere to work?
Imagine asking for a blue travel mug that can go in the dishwasher. A useful answer might recommend a few mugs. An agent can go further: open a shop, read the product details, choose an option and work towards checkout. To understand how, we can follow what happens between your request and the next click.
A computer somewhere else
The word “cloud” describes where the work happens. The browser can run on a remote server, a physical computer reached over the internet. Your phone is where you send the request; the shopping takes place on that other machine.
Giving each person a computer sounds expensive. But a server can be divided into several virtual machines: separate computers created in software, each with its own operating system. They share the physical machine’s resources. The picture below shows how several people can have separate workspaces on one server.

The AI model can live elsewhere again. The browser’s computer sends it information and receives instructions in return. Software on the computer carries out those instructions, then reports what happened. This connection lets a model do something with its answer, such as opening a page or typing into a search box.
There is also a practical reason to keep a workspace around. A browser can retain a signed-in session, and the agent’s software can save a record of its work. After an interruption, that record can help it pick up where it left off.
How words become clicks
Suppose the mug’s product page is open. In one kind of system, the model receives a screenshot. It can examine the image and request a click at a particular spot. A small piece of software turns that request into an actual browser action, just as a mouse click would.
Other systems give the model a structured description of the page. That description might identify a heading, some text and a button labelled “Add to cart.” This is called an accessibility snapshot. It gives the model another way to work out which control it can use.

After the click, the model gets another look at the page. Did the cart appear? Is the mug there? If nothing changed, the agent cannot safely assume the click worked. Each result becomes information for the next decision. This repeating cycle of observing, choosing an action and checking the result is how a browser agent works through a task.
That cycle also gives it room to change course. The first blue mug might say “hand-wash only.” The agent could return to the results and open another product. The goal stays the same, while the route changes with what it finds. Researchers have explored this combination of reasoning and action in a method called ReAct. It is one approach to building such a loop.
Where the skill comes from
Being given a browser does not teach a model how to use it. Computer-use training helps connect what appears on a screen with actions. Anthropic’s early work, for example, trained a model to interpret screenshots and locate where to move a cursor. Recognising a button and choosing the right place to click are abilities the system has to learn.
Researchers can also create places to practise. WebShop is a simulated shopping website where an agent searches for products that match a written request. Its researchers trained agents using human demonstrations and feedback about their choices. A selection could be judged against the requested product features, options and price. Learning from that feedback is called reinforcement learning.
This helps explain how a shopping agent can improve at choosing products. It does not mean a good score in a simulated shop guarantees a successful order on a real website. The real task also involves access to accounts and permission to act.
Getting to checkout is only part of the job
An agent can be allowed to read information without being allowed to send a message or spend money. Those permissions can be checked by software outside the agent making the decisions. That matters because a website can contain instructions designed to mislead the AI, an attack called prompt injection. A shop’s page should not get to decide what the agent is allowed to do with your account.
An account’s sign-in details can also be kept in separate storage. Software uses them when needed, so the model can request an action without being given the secret itself. This is one way to connect an agent to a service while limiting what it can see.
Back at the mug shop, reaching checkout would show that the agent had managed a sequence of smaller jobs: reading, choosing, clicking and checking. If the task was to prepare the order and ask before buying, that is where it should wait. The next decision belongs to the person paying for the mug.
