🖱️ Fara1.5-9B — Computer Use Agent (Visual Grounding)

Fara1.5-9B by Microsoft Research AI Frontiers is a 9B vision-language computer use agent that predicts pixel-level actions from browser screenshots.

How it works: Upload a screenshot of a web page and describe the task. Fara analyzes the screenshot and predicts the next action — a click at a specific pixel coordinate, text to type, a scroll, etc. — visualized with a marker on the image.

The model grounds click/drag targets in a normalized 0–1000 coordinate space, mapped onto a 1440×900 viewport. Coordinates are rescaled to match the actual screenshot dimensions for the overlay.

Upload a screenshot and describe a task to see the predicted action.

Try an example — real screenshots with grounded actions

Examples