---
source: 'https://howaiworks.ai/blog/markov-ai-computer-use-large-dataset'
section: blog
title: 'Computer Use Large: The Largest Open-Source Dataset for AI Agents'
description: >-
  Markov AI releases 'computer-use-large', a massive dataset of 48,000+ screen
  recordings for training AI agents to use professional software.
date: '2026-03-16'
author: HowAIWorks Team
tags:
  - Markov AI
  - Computer Use
  - AI Agents
  - Open Source Dataset
  - HuggingFace
  - GUIs
  - Machine Learning
  - Dataset
readingTime: 4 minutes
---

# Computer Use Large: The Largest Open-Source Dataset for AI Agents

> Markov AI releases 'computer-use-large', a massive dataset of 48,000+ screen recordings for training AI agents to use professional software.

## Introduction

The race to build reliable **Computer Use Agents**—AI models that can navigate digital interfaces just as humans do—has hit a significant milestone. Markov AI has just released **Computer Use Large**, the largest open-source dataset of professional computer work ever made available. 

As AI agents move from simple chat interfaces to directly controlling desktop software, the need for high-quality, diverse demonstration data has never been greater. This dataset provides the necessary "trajectories" for models to learn how professional software is actually used in the real world.

## A Massive Scale for Professional Workflows

The scale of **Computer Use Large** is unprecedented in the open-source community. Hosted on HuggingFace, the dataset comprises:

- **48,478 screen recordings**: High-quality captures of software interfaces.
- **~12,300 hours**: Over a year of continuous usage if watched back-to-back.
- **Trimmed & Focused**: Every video has been stripped of audio and non-essential content like intros, talking heads, and transitions to keep the focus entirely on the GUI interaction.

## Diverse Software Categories

One of the most valuable aspects of this release is its focus on "professional" software. Unlike general web-browsing datasets, this collection includes sophisticated desktop applications where precise control is required:

- **AutoCAD**: Engineering and design workflows.
- **Blender**: 3D modeling and animation.
- **Excel**: Complex spreadsheet manipulations and data entry.
- **Photoshop**: Professional image editing and graphic design.
- **Salesforce**: Enterprise CRM workflows and industrial data management.
- **VS Code**: Real-world programming and development environments.

The data is organized by category, making it easy for researchers to fine-tune models on specific software domains.

## Built for Agentic Training

This dataset isn't just a collection of videos; it's a foundation for the next generation of **Large Action Models (LAMs)**. By providing thousands of examples of how humans navigate nested menus, use keyboard shortcuts, and interact with complex UI elements, Computer Use Large enables:

1. **Behavioral Cloning**: Teaching agents to mimic expert workflows.
2. **GUI Understanding**: Improving the model's ability to "see" and interpret UI components.
3. **Workflow Evaluation**: Providing a robust benchmark to test how well agents can follow multi-step professional tasks.

## Conclusion

The release of **Computer Use Large** by Markov AI is a major win for the open-source AI ecosystem. By "democratizing" access to high-quality professional trajectories, Markov AI is lowering the barrier for developers and researchers to build agents that aren't just clever conversationalists, but capable digital workers. As we move closer to a world where AI can assist with complex CAD designs or manage enterprise sales pipelines, datasets like this will be the fuel that drives that transformation.

## Sources

- [Markov AI Computer Use Large on HuggingFace](https://huggingface.co/datasets/markov-ai/computer-use-large)
- [Dataset Metadata and Organization](https://huggingface.co/datasets/markov-ai/computer-use-large#data-organization)

## Frequently Asked Questions

### What is the 'Computer Use Large' dataset?

It is a large-scale open-source dataset by Markov AI containing 48,478 screen recordings (~12,300 hours) of professional software usage for training AI agents.

### What software categories are included in the dataset?

The dataset covers several professional applications including AutoCAD, Blender, Excel, Photoshop, Salesforce, and VS Code.

### How was the data processed?

All videos were trimmed to remove intros, outros, and transitions, and audio was stripped to focus purely on the screen recording content.

### What is the license for this dataset?

The dataset is released under the CC-BY-4.0 license, allowing for broad use and adaptation with proper attribution.

---

Source: https://howaiworks.ai/blog/markov-ai-computer-use-large-dataset — HowAIWorks.ai
