The promise of general-purpose robots in our homes, powered by sophisticated Robot Foundation Models (RFMs), hinges on more than just sheer capability. A new study reveals that for these AI systems to be truly useful and safe, users need a deeper, more nuanced understanding of their performance, extending far beyond simple task success rates.

Beyond the Bottom Line: Unpacking Performance Metrics

Robot Foundation Models are designed to be adaptable, tackling tasks they weren't explicitly trained for. This adaptability, however, introduces inherent risks. When a user asks an RFM-powered robot to perform a novel task, failure can be costly. Therefore, ensuring users grasp the robot's limitations and potential failure points is paramount. Researchers at [Redacted University, affiliation implied by research context] explored how non-experts interpret performance data from these advanced AI systems.

Their study, published on arXiv (arXiv:2602.03920v1), involved participants viewing evaluation data from various published RFM projects. This data included the typical Task Success Rate (TSR), descriptions of why tasks failed, and video evidence. The findings suggest that while users do understand TSR as intended by experts, they also place significant value on qualitative information.

“While TSR is intuitive to experts, it is necessary to validate whether novices also use this information as intended,” the paper states. The research team found that non-expert users not only interpreted TSR correctly but also highly valued supplementary information, such as detailed failure case descriptions, which are often omitted from standard RFM evaluations.

The Human Element: Context and Confidence

This indicates a critical gap in how RFM performance is currently communicated. Experts might be satisfied with a numerical success rate, but everyday users require richer context to make informed decisions about deploying robots in their homes. Understanding why a robot fails, not just how often, builds trust and allows for more responsible interaction.

Furthermore, the study highlighted a strong user desire for both retrospective and prospective performance information. Participants wanted to see real data from past evaluations to gauge the RFM's historical reliability. Crucially, they also wanted the robot itself to provide an estimate of its likely performance on a new, untested task.

This dual requirement—understanding past performance and predicting future success—underscores the need for RFMs to not only execute tasks but also to self-assess and communicate their confidence levels. Such a feature would empower users to better judge whether to proceed with a given command, thereby mitigating risks and fostering a more collaborative human-robot relationship.

"Furthermore, we find that users want access to both real data from previous evaluations of the RFM and estimates from the robot about how well it will do on a novel task."

— arXiv:2602.03920v1

The implications of these findings are substantial for the development and deployment of domestic robotics. As RFMs become more prevalent, clear, and comprehensive performance communication will be as important as the underlying AI technology itself. Developers must move beyond simplistic metrics to build systems that can transparently convey their capabilities and limitations, ensuring that users can confidently integrate these powerful new tools into their lives.