A sample dataset of computer-use tasks for professional software testing.
Why it made the shelf
It provides a dataset for computer-use tasks which is used for training and testing AI agents.
Seen on
Best matches
Public trail
1 reference from 1 publisher.
3 tools
An open benchmark for evaluating AI performance on system-on-chip tasks.
An agentic mobile app penetration testing tool.
Perform pass at k reliability testing for AI coding skills in Claude Code and Codex.