The Demo Is No Longer Enough
A humanoid walks across a stage, turns, picks up a box and places it on a table. The sequence lasts perhaps half a minute. It is smooth enough to attract cameras and simple enough to communicate instantly. For the last few years, much of the public discussion around humanoid robotics has been built on scenes like this.
Industrial engineering starts where the applause stops.
The first questions are not theatrical. How fast was the robot actually moving? How much load could it handle? What happened when the surface changed? Could it repeat the task fifty times? What happened as the actuators warmed up? Was the sequence autonomous, scripted or teleoperated? How did the machine react when communication was interrupted? How much energy did a successful cycle consume? And if the robot did something unexpected, how quickly could it be brought into a safe state?
These questions are beginning to move from conference corridors into laboratories. In 2026, Fraunhofer IPA introduced a modular benchmark intended to evaluate humanoid robots against industrially relevant criteria. The significance is larger than another set of performance numbers. It represents a change in the way the technology is being discussed: away from isolated demonstrations and toward measurable operational suitability.
That shift is overdue. A robot can be impressive without being useful. It can be useful without being dependable. And it can be dependable at one task while being economically inferior to a mobile manipulator, an AMR combined with a cobot, fixed automation or a human worker.
All of those statements can be true at the same time.
The deeper reason this matters is procurement risk. A humanoid is not bought as a collection of actuators and algorithms; it is bought as a promise that a physical job can be transferred from one operating model to another. That promise has several failure modes. The robot may lack a capability entirely. It may possess the capability but only under narrow conditions. It may perform adequately but need more human supervision than expected. Or it may work technically while its charging, maintenance and recovery demands undermine the business case.
A benchmark becomes valuable when it separates those failure modes instead of averaging them into one attractive score. For an industrial user, a machine that is excellent at locomotion but weak in safe contact is not “almost ready” if the application is collaborative. A robot with strong manipulation but poor cleanability is not a near miss in a contamination-sensitive process. The missing capability can dominate the whole decision.
A serious benchmark therefore has to do more than expose capability. It has to create a common language for comparing machines whose architectures, software stacks and operating concepts may be radically different. Fraunhofer IPA has made a strong start. The next engineering problem is to define the layer that comes after capability: the mission itself.
What Fraunhofer IPA Gets Right
Fraunhofer IPA divides its humanoid benchmark into six areas: technologies and basic capabilities, complex capabilities, cleanroom suitability, functional safety, cybersecurity and energy efficiency. The structure is deliberately modular. Where possible, the institute links tests to established industrial standards, including ISO 14644 for cleanroom topics and ISO 10218 and ISO/TS 15066 for safety-related measurements.
That breadth matters. A humanoid entering a factory is not only a locomotion system with two arms. It is a moving electrical, mechanical and computational platform that may operate close to people, connect to networks, enter sensitive production environments and carry enough stored energy to sustain useful work for hours.
The Unitree G1 evaluation published by Fraunhofer IPA illustrates the value of looking beyond headline specifications. The tested configuration was a G1 EDU-4 with Dex3-1 three-finger hands and firmware version 1.04. Fraunhofer reported measured walking speeds of approximately 0.49 m/s in a slower mode and 0.84 m/s in a faster mode. It also reported that horizontally extending the arms could lead to thermal shutdown of the actuators after roughly one to two minutes even without an additional payload. Those two observations belong to very different brochures. One describes mobility. The other describes duty cycle.
The safety tests were equally revealing. Fraunhofer reported collision forces above 500 N in some whole-body and fast-arm movements and identified practical concerns including the lack of an emergency-stop button directly on the robot and pinch-point risks around joints. The cybersecurity work found a documented remote-code-execution vulnerability in the Bluetooth interface of the software version examined; Fraunhofer states that the issue has since been closed by a later update.
Cleanroom testing produced a different kind of result. Particle emission and outgassing were promising enough for Fraunhofer to describe potential suitability for ISO Class 5 environments, while hygienic design and cleanability remained problematic because of inaccessible gaps and joint geometries. This is a useful reminder that “cleanroom suitable” is not a single property. Particle behavior, outgassing, surface design and maintainability can point in different directions.
Energy testing adds another layer. Fraunhofer measured roughly 154 W while standing, around 272 W walking on level ground, and about 283 W at a 10 percent incline. For its defined one-hour standard scenario, the average was approximately 239 W. The institute derived operating times of 2 hours 49 minutes for standing and 1 hour 49 minutes for a typical scenario combining standing and walking.
These numbers are valuable precisely because they are not spectacular. They are the kind of numbers an engineer can put into a spreadsheet and begin asking better questions.
There is also a systems-engineering reason to keep the individual measurements visible. Composite scores are convenient for rankings but dangerous for design. If battery runtime, collision behavior, grasping performance and cybersecurity are collapsed into a single index, the number may help a purchasing discussion while hiding the mechanism that engineers need to fix. Industrial maturity is rarely one-dimensional. It is the intersection of several constraints, and the limiting constraint can change from one application to the next.
That is the strength of a modular benchmark. A semiconductor fab, a warehouse and a machine shop do not need the same robot. Cleanroom emissions can dominate one decision, while payload, autonomy or recovery behavior dominates another. The evaluation should preserve enough resolution to ask not only “Which robot scored higher?” but “Which limitation would stop deployment in my process?”
The benchmark also has an important strategic ambition: Fraunhofer says the results should make humanoids comparable not only with one another but also with established automation components. That is where the exercise becomes industrial rather than merely robotic.
The two questions are complementary rather than competing: qualification establishes what the machine can do; a mission profile defines what the architecture must sustain while doing useful work.

Power Is Not Productivity
The energy measurements expose a deeper problem in humanoid evaluation. Average power consumption is necessary information, but it is not yet a productivity metric.
Consider two hypothetical machines. Robot A consumes 250 W while executing a task and completes forty successful cycles per hour. Robot B consumes 400 W but completes one hundred and twenty successful cycles in the same period. Looking only at instantaneous power, Robot A appears more efficient. Looking at energy per completed task, Robot B may be substantially better.
The difference is not academic. Factories do not purchase watts. They purchase output under constraints.
Useful metrics therefore have to connect energy with work: watt-hours per successful pick, watt-hours per completed assembly operation, tasks per battery charge, or energy per kilogram moved through a defined distance. The exact metric depends on the application, and careless definitions could easily become misleading. But the principle is robust: consumption has to be normalized against useful output.
The same is true of speed. A walking velocity of 0.84 m/s is a clear and repeatable physical result. It says little, by itself, about transport productivity. A robot that walks quickly but spends thirty seconds aligning itself at every workstation may be slower over a mission than a machine with lower peak speed and better positioning behavior. A highly dexterous hand may be irrelevant if the robot requires frequent operator intervention. A large nominal payload may be economically meaningless if carrying it triggers thermal derating after a short interval.
Task quality adds another complication. A pick is not successful merely because an object leaves one surface and reaches another. Orientation may matter. Placement tolerance may matter. Damage may matter. Cycle time may matter. A robot that completes ninety-nine percent of cycles but occasionally drops a fragile component can be unacceptable, while a slower robot with predictable behavior may be preferred. “Successful task” therefore has to be defined with the same care as power or speed.
The most useful normalization may eventually be application-specific. Logistics could focus on energy per correctly transported tote or per kilogram-meter. Machine tending could focus on successful machine cycles per charge and interventions per shift. Assembly might emphasize completed operations within force and position tolerances. The common principle is not one universal unit. It is that resource consumption should be related to a verified outcome.
Benchmark numbers begin to behave like clues. Each describes a real property, but the operational conclusion appears only when the clues are connected.
For industrial humanoids, that connection is the mission.
The Missing Mission Layer
Automotive engineering has lived with this problem for decades. A vehicle cannot be characterized adequately by measuring its engine at one operating point. Real use consists of acceleration, cruising, braking, waiting, gradients, traffic and temperature. Standardized drive cycles are imperfect abstractions of reality, but they allow engineers to compare systems under a shared workload.
Humanoid robotics needs an equivalent idea, although not necessarily one universal cycle.
A reference mission profile would define a repeatable time-domain sequence such as: idle, observe, walk, accelerate, stop, reach, grasp, lift, carry, place, interact, wait and repeat. Instead of asking only whether each capability exists, the profile asks what happens when those capabilities are coupled over time.
The distinction is important. Walking for sixty seconds from a cold start is not the same as walking after repeated lifting. Holding an arm horizontally for a laboratory test is not the same as reaching hundreds of times during a shift. A battery measured at a comfortable state of charge may behave differently near its operational limits. Compute and perception loads rise when scenes become dynamic. Communications traffic changes when the robot moves between network zones or coordinates with a fleet. Thermal conditions accumulate quietly until a limit appears.
The profile should also contain transition states. This sounds like a small detail, but transitions often create the highest stresses. Starting to walk can demand more peak power than steady walking. Catching balance after a manipulation error can produce sharp actuator loads. Switching from navigation to fine manipulation can shift compute demand from mapping and motion planning toward perception and control. A robot that looks efficient in steady states may be inefficient in the transitions that dominate a short-cycle application.
Recovery deserves its own place in the profile as well. Real missions contain imperfect grasps, obstructed paths, displaced objects and uncertain detections. If recovery is excluded, the profile describes an idealized robot living in an idealized factory. Including controlled disturbances makes the workload less elegant but more representative. It also provides a way to measure how much extra energy, time and computation are consumed when the world refuses to cooperate.
A useful mission profile therefore needs more than one trace. It needs a state vector. At each point in time, the profile can describe locomotion demand, manipulation demand, payload, perception intensity, compute state, communication activity, safety context and environment. The resulting mission is not a robot specification. It is a workload definition.
The analogy with WLTP is useful but incomplete. Passenger vehicles share a relatively narrow core function. Humanoids may inspect equipment, move totes, load machines, handle tools, assemble parts or perform logistics tasks in environments with very different requirements. One universal mission could become artificial very quickly.
A better approach may be a family of reference missions: logistics, machine tending, inspection, material handling, assembly and perhaps service. Each would preserve a common measurement framework while changing the ratio of locomotion, manipulation, payload and interaction.
That still leaves difficult questions. How long should a mission run? How many repetitions are statistically meaningful? How should task quality be scored? How should human assistance be counted? Which signals are available from a closed commercial robot, and which require internal telemetry from the manufacturer?
Those questions define the work required to make the concept useful rather than weaken the case for it.
Once the mission states share one time axis, their different subsystem consequences become visible: mechanical demand can peak in seconds, while thermal load accumulates more slowly; sensing may remain continuously active while communication arrives in bursts.

What the Architecture Must Survive
Once a common mission exists, the engineering picture changes. The same five-minute or one-hour sequence can become the input to multiple subsystem models.
Battery engineers can observe current peaks, average power, regenerative events, state-of-charge movement and charging requirements. Motor-control engineers can examine joint torque, speed, RMS current and inverter utilization. Thermal engineers can identify whether repeated manipulation creates local hotspots long before the battery is empty. Compute teams can map perception, planning and control workload against task state. Network architects can see whether information flow is constant or bursty and whether safety-critical communication competes with high-bandwidth perception data.
That common workload is more powerful than a collection of disconnected assumptions.
Without it, the battery may be sized against one imagined duty cycle, the actuators against another and the communications system against a third. Each subsystem can be locally optimized while the complete robot remains poorly balanced.
With it, architecture questions become testable. Is the robot energy-limited or power-limited? Which joints dominate thermal stress? How much energy can realistically be recovered? Does a higher-voltage backbone reduce current enough to matter at the mission level? Does a more efficient actuator save battery capacity, reduce cooling demand, lower cable mass and then indirectly reduce locomotion energy? Does additional compute improve perception enough to reduce failed grasps and therefore reduce energy per successful task?
The answers may cascade across the machine.
For semiconductor and component design, the distinction matters especially because component stress is not defined by average robot power. Power electronics age through electrical and thermal cycles. Batteries experience charge throughput and peak current. Motor drivers see duty-dependent losses. Communication devices encounter changing traffic and fault conditions. Sensors may switch between high-performance and low-power states. A mission profile provides the time axis on which these interactions become visible.
A common time base also exposes design trade-offs that are otherwise discussed in isolation. A larger battery may extend runtime but increase robot mass, which increases locomotion energy and joint loading. More powerful compute may improve perception and planning but add electrical and thermal demand. Higher actuator margins may improve peak performance but penalize mass and efficiency if they are rarely used. Better sensing can reduce task failures, yet it can also increase bandwidth and compute requirements. The optimum sits at system level, not inside any one component.
Mission profiles become especially valuable before a design is frozen. They can be simulated long before a physical robot exists. Different architectures can be run against the same workload and compared on peak power, energy, thermal margin, estimated mass and expected task throughput. Later, measured data from prototypes can replace assumptions without changing the mission definition. The profile becomes a bridge between simulation, laboratory testing and field validation.
It could also support lifetime engineering. Repetition counts transform a five-minute workload into thousands of thermal cycles, mechanical reversals and battery events over months of service. The extrapolation must be done carefully, but at least the stress history begins from a common operational hypothesis rather than arbitrary component-level duty factors.
The architecture is no longer evaluated against a brochure maximum. It is evaluated against a workload it has to survive repeatedly.
Dependability Changes the Score
There is another variable that can overturn almost every ranking: repeatability.
A robot that completes a task once has demonstrated capability. A robot that completes it ten thousand times with a known distribution of failures has demonstrated something closer to industrial performance.
NIST is moving in this direction with its proposed Humanoid Robot Baseline Performance Benchmark. The proposal is built around quantifiable locomotion and manipulation tasks and explicitly includes coordinated loco-manipulation, whole-body awareness and minimal reasoning. NIST is also seeking industry input on which tasks are sufficient to establish a meaningful baseline.
The NIST effort is complementary to Fraunhofer IPA rather than a substitute for it. Together they show a field that is still deciding what “good” means.
Dependability requires metrics that are less attractive in a demonstration video but far more important in a factory: successful task percentage, grasp failures per thousand attempts, falls, autonomous recoveries, human interventions, unexpected resets, localization losses, thermal derating events and mean time to restore operation after a fault.
Long-duration reliability measures such as MTBF are difficult to obtain in short laboratory campaigns, and small test samples can create false confidence. That limitation should remain visible. A benchmark should not pretend that fifty repetitions reveal lifetime reliability. But even a modest number of repetitions can expose the difference between deterministic success and probabilistic success.
Autonomy needs similar discipline. A teleoperated robot and an autonomous robot may perform the same physical movement while representing completely different industrial systems. Scripted execution, task-level autonomy, environment adaptation, instruction-driven behavior and self-recovery should not be collapsed into one capability score.
The hidden variable is human support. How many minutes of engineering are required to teach the task? How often does an operator intervene? What happens when an object is displaced by ten centimeters? Does the system recover from an incomplete grasp, or does a person reset the sequence?
These measures can be uncomfortable because they reveal the operational machinery behind the demonstration. That is exactly why they matter.
Human intervention is especially important because it can hide inside otherwise excellent statistics. A robot may show high task success if an operator quietly resolves ambiguous situations, repositions objects or restarts failed sequences. From a factory perspective, those interventions are part of the cycle cost. They consume labor, interrupt flow and complicate scaling from one robot to a fleet.
A useful benchmark should therefore record not only whether intervention occurred but why. Was the cause perception uncertainty, manipulation failure, localization loss, safety logic, communication, planning or an unknown software state? The categories turn a frustrating event into engineering evidence. Over time, they also reveal whether improvements are genuinely increasing autonomy or merely shifting operator effort elsewhere.
Fleet operation raises the stakes further. A one-percent intervention rate may look tolerable on a single machine running occasional tasks. Across one hundred robots executing thousands of cycles, the same rate can create a permanent support workload. Reliability and autonomy metrics have to be interpreted at deployment scale.
Safety is beginning to confront the same systems problem. ISO/CD 25785-1, currently under development, addresses safety requirements for industrial mobile robots whose stability depends on active control, explicitly including bipedal and other dynamically stable forms. The standard is not yet a published International Standard, and its eventual requirements should not be anticipated as settled. Its existence nevertheless signals that humanoid-like mobility is entering a regulatory space where dynamic stability, failure behavior and integration can no longer be handled by analogy alone.
A robot that loses a camera, a network link, localization confidence or battery margin still has to do something. The interesting question is not whether a fault occurs. Faults are inevitable. The question is whether the machine falls, freezes, enters a safe hold, degrades gracefully, returns to a known pose or requests assistance.
Industrial trust is built from that behavior.
Two Standards, Not One
There is a temptation in emerging technologies to search for one grand benchmark. Humanoid robotics will probably resist that.
Fraunhofer IPA’s Humanoid Capabilities Navigator already separates a different problem from physical testing. It classifies robots and applications across five maturity levels and four capability areas: mobility, manipulation, cognition, and safety and security. In practical terms, the Navigator helps frame what level of capability an application requires.
The IPA benchmark then asks what the machine actually delivers across a broader set of measurable industrial criteria.
A reference mission profile would add a third question: what workload does the selected application impose over time?
Those three layers should not be collapsed.
The first is requirements. The second is validation. The third is duty cycle.
Below them sits subsystem stress: battery, power electronics, actuation, thermal behavior, sensing, compute and communications. Below that sits lifecycle economics: energy per useful task, maintenance, uptime, component life and ultimately cost per successful work cycle.
A layered view avoids a false competition between benchmarking approaches. Fraunhofer’s work does not need to be replaced by another benchmark that measures the same things slightly differently. Its strength is precisely that it is creating objective evidence around industrial suitability. A mission profile should complement that evidence by connecting capabilities to time, repetition and architecture.
There is a practical advantage to this separation. It allows different institutions and companies to contribute without waiting for one organization to own the entire stack. A research institute can improve capability tests. Standards bodies can define safety requirements. Robot manufacturers can publish mission telemetry. Component suppliers can build simulation models. End users can define representative workloads. The interfaces between those layers matter more than institutional ownership.
For the market, this could reduce a recurring source of confusion: comparing results produced under different assumptions. A robot advertised with a long runtime may have been tested mostly standing. Another may report a shorter runtime under continuous walking. A third may quote battery capacity without a defined workload at all. A reference mission does not make the machines identical; it makes the assumptions visible.
The automotive analogy remains useful here. Euro NCAP and WLTP answer different questions. One evaluates safety-related performance under defined tests; the other creates a standardized operating cycle for consumption and emissions comparison. Neither describes the entire vehicle, and neither eliminates engineering judgment. Together they create a much richer evidence base than either could provide alone.
Humanoid robotics is more diverse and more immature, so the mapping cannot be literal. But the principle survives: qualification and mission are different abstractions.
The Metric That Matters: Useful Work
The most important comparison in humanoid robotics may eventually be the one that does not compare humanoids with humanoids.
A factory manager rarely begins with the requirement “buy a biped.” The requirement is closer to “move these materials,” “tend this machine,” “inspect this area,” or “perform this assembly operation under these constraints.” The alternatives may include a humanoid, a mobile manipulator, an AMR plus cobot, fixed automation, process redesign or continued human work.
The endpoint of benchmarking changes with that procurement question.
Walking speed, payload, battery runtime and gripping force remain necessary. But they become intermediate variables. The industrial outcome is useful work delivered safely and repeatedly at an acceptable cost.
A mature evaluation framework could therefore converge on measures such as energy per successful task, successful tasks per charge, interventions per hundred tasks, recoveries per failure, thermal deratings per thousand cycles and cost per completed work unit. None of these is a universal standard today, and each needs careful definition. A bad metric can distort engineering as easily as a bad benchmark.
The direction is nevertheless becoming clearer.
Fraunhofer IPA has already made an important move by forcing humanoid robotics to confront evidence beyond the stage demonstration. Its tests expose thermal limitations, safety behavior, cybersecurity, cleanliness and energy consumption that marketing material can obscure. NIST is developing another baseline around repeatable physical capabilities. ISO is working on safety requirements for dynamically stable industrial mobile robots. The pieces of a more disciplined industrial framework are beginning to appear.
The missing bridge is the mission.
Define what the robot has to do over time. Run the machine through that workload. Measure not only whether it can finish, but how often, how autonomously, with what energy, under what thermal stress, with how many interventions and with what failure behavior.
Then connect the result back to the architecture.
That is when a benchmark stops being a scorecard and becomes an engineering instrument.
The decisive question for industrial humanoids is no longer whether a machine can walk across a stage and pick up a box. It is whether the complete system can convert stored energy, computation, sensing and motion into dependable useful work, hour after hour, under conditions it did not choose.
Benchmark the robot.
Then standardize the mission.
