Human Barista or Embodied AI Barista—which would you choose?

The same cup of coffee.
One made by an experienced human barista. The other by an embodied robot that takes no breaks, never loses focus, and can keep working cup after cup.
Which would you choose?
How Does a Human Barista Make a Great Cup of Coffee?
Remove the portafilter, dose the grounds, distribute and tamp them, then lock the portafilter into place. Position the cup and start the extraction. Once the coffee is ready, serve it, empty the used grounds, rinse the portafilter, and return the tools to their places.
It may take only a few minutes, but behind it lies a constant stream of observations and decisions:
Is the cup positioned correctly? Is the portafilter aligned? Are the grounds properly tamped? Have the grounds from the previous cup been cleared away?
A skilled barista can make all of this look effortless.
But what happens when they have to make 100 cups in a row?
Their hands grow sore, fatigue sets in, and concentration begins to slip. After all, people are not machines.
So we decided to teach this entire coffee-making process to a new “colleague” that never gets tired.
How Do You Teach a Robot to Make Coffee?
The answer is not a lengthy instruction manual for a robotic arm, nor a set of fixed coordinates for every movement.
It is more like giving the robot hands-on barista training.
Step one: collect the data.
A human first demonstrates the operations using a real coffee machine and tools. The system simultaneously records what the robot “sees” and how it “should move”: camera images, task instructions, the robotic arm’s position, motion trajectories, and grasping and placement actions all become training data.
Each training example captures a complete demonstration or part of one. From these demonstrations, the robot gradually learns how to grasp a portafilter when it sees one, how to align it with the coffee machine, and which step should follow extraction.
Step two: train the model.
We fine-tune our in-house VLA model using a small amount of real-world demonstration data. VLA stands for Vision, Language, and Action. At its core, the model addresses three questions:
“What am I seeing?”
“What task do I need to complete now?”
“How should I move next?”
Through training, the model learns to connect visual input, tasks, and actions. Instead of simply following fixed coordinates, the robot predicts its next action based on the environment in front of it.
Step three: put it to the test on the real robot.
Finishing training does not instantly turn the robot into a barista. We still need to run repeated trials on real equipment.
What if the cup is slightly out of place? What if the portafilter has moved? What if a step has not been fully completed?
Each failure reveals new problems. We then collect additional data for those situations, adjust the training material, and fine-tune the model again, gradually teaching the robot to handle changes in the real world.
The process can be summed up as:
Human demonstration → Data collection → Model training → Real-world robot testing → Additional data collection → Retraining.
Through repeated iteration, a single model eventually links the separate actions of grasping, moving, dosing, tamping, extracting, serving, and cleaning into one complete, long-horizon task.
Getting a robot to perform a single action is nothing unusual.
The real challenge is enabling it to understand which step it has reached, what comes next, and how to execute the entire process reliably from start to finish.

On Cup 101, It Still Does Not Get Tired
At WRC 2026, this embodied robot barista served more than 100 cups a day on average and made over 400 cups in total. Among the coffee-making projects at the event, it ran continuously for the longest time and produced the most cups.
It takes about as long to make a cup as a human barista does.
But after making 100 cups, a human may just want to sit down and rest. After making 100 cups, the robot simply gets started on cup 101.
One visitor commented:
“The other coffee robots aren’t moving. Yours is the only one that keeps making coffee, and it does it just like a person.”
So who wins: the human barista or the embodied robot barista?
If you want conversation, creativity, and a cup with a personal touch, human baristas remain irreplaceable.
But if you need consistent service, cup after cup, without fatigue, distraction, or interruption, embodied robot baristas are showing what else is possible.
It may not be faster than a person, but it can carry out the entire process of making a good cup of coffee, just as a person would.
And then keep going.
Now, with these two cups on the counter—would you like to try the one made by the embodied AI barista?
