# Deskwright (pypi · deskwright)

Agents use real GNOME/Wayland apps: on your screen, or an invisible second desktop you never see

- Trust score: 89/100 (high trust)
- Change this week: +17
- Registry status: active
- Liveness: live
- Owner verified: no
- Last scored: 2026-09-28

## Components

- pypi · `deskwright`: 89/100 (this document), [markdown](https://verifymcp.io/servers/tristanmuzzu-deskwright/deskwright.md), [page](https://verifymcp.io/servers/tristanmuzzu-deskwright/deskwright)

## Channel facts

- Registry: `pypi`
- Package: `deskwright`
- Version: `0.1.1`
- Transport: `stdio`

## Trust breakdown

How this component scores in each security and reliability category. Every signal is checked automatically from public evidence about the published package, including repeated runs of it in an isolated sandbox, and we only credit what we can confirm. Scores are 0–100 per category. Scoring method: https://verifymcp.io/docs/scoring (what has changed: https://verifymcp.io/docs/scoring/changelog)

Scored 2026-09-28.

- **Supply Chain Security**: 100/100
  - No malware found by supply-chain analysis.
  - No known CVEs affecting this package version or its production dependencies.
  - Runs setuptools.build_meta at install time, a recognised build step with no custom scripting around it.
  - 0 of 1 dependencies flagged as unhealthy.
- **Provenance & Transparency**: 100/100
  - Source repository is publicly reachable at the declared URL.
  - Cryptographically verified build provenance (signed, bound to tristanmuzzu/deskwright).
  - Clear OSI-approved license (Apache-2.0).
  - Actively maintained (last published 26 days ago).
  - Publishes a security disclosure policy (SECURITY.md).
- **Schema Quality & AI Usability**: 66/100
  - AI-judged instruction clarity (excellent).
  - Context-footprint check failed: tool/resource definitions use about 7389 tokens (~223/item across 33 items; 33 tools + 0 resources), over budget; trim descriptions and params.
  - Usage-examples check failed: none of the tools include examples.
- **Stability & Change Management**: 90/100
  - Stability observed for 27 of 30 days with no destabilising changes; credit accrues until the full window elapses.
- **Tool Coverage**: 90/100
  - 100% of tools have a non-trivial description (not blank, and not just the tool's name).
  - 71% of tool parameters carry a description.
- **Tool Safety**: 75/100
  - No prompt-injection markers were found in the server instructions, tool names or descriptions we captured.
  - 0 of 1 tool(s) whose name or description implies an irreversible operation declare an MCP destructiveHint annotation; "press_keys" implies "send" and declares no destructiveHint at all, which the MCP spec reads as destructive by default.
  - An AI judge read all 33 captured unit(s) of tool text and found none that tries to manipulate the model reading it.
- **Capabilities**: 60/100
  - Spec-recency check failed: implements MCP spec 2025-06-18; the latest is 2026-07-28.

## Install

### How do I install the Deskwright MCP server?

Deskwright runs locally as a PyPI package, launched with uvx deskwright. Ready-made configuration for Claude, Cursor, VS Code, Codex and 5 more is on this page, copied from each client's own documentation.

### Claude

```bash
claude mcp add tristanmuzzu-deskwright -- uvx deskwright
```

### Cursor

```json
{
  "mcpServers": {
    "tristanmuzzu-deskwright": {
      "command": "uvx",
      "args": [
        "deskwright"
      ]
    }
  }
}
```

### VS Code

```json
{
  "servers": {
    "tristanmuzzu-deskwright": {
      "command": "uvx",
      "args": [
        "deskwright"
      ]
    }
  }
}
```

### Codex

```bash
codex mcp add tristanmuzzu-deskwright -- uvx deskwright
```

### opencode

```json
{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "tristanmuzzu-deskwright": {
      "type": "local",
      "command": [
        "uvx",
        "deskwright"
      ],
      "enabled": true
    }
  }
}
```

### OpenClaw

```bash
openclaw mcp add tristanmuzzu-deskwright --command uvx --arg deskwright
```

### Hermes

```yaml
mcp_servers:
  tristanmuzzu-deskwright:
    command: "uvx"
    args: ["deskwright"]
```

### Netclaw

```json
{
  "McpServers": {
    "tristanmuzzu-deskwright": {
      "Transport": "stdio",
      "Command": "uvx",
      "Arguments": [
        "deskwright"
      ]
    }
  }
}
```

### Vellum

```bash
assistant mcp add tristanmuzzu-deskwright -t stdio -c uvx -a deskwright
```

### Other

```json
{
  "mcpServers": {
    "tristanmuzzu-deskwright": {
      "command": "uvx",
      "args": [
        "deskwright"
      ]
    }
  }
}
```

## Changelog

Every change recorded for this component, newest first. Days that predate change tracking, or that we cannot explain, say so: "we were watching and nothing happened" and "we were not watching" are different claims.

### 2026-09-28 (score 89, +11)

- [functional] We updated how we score, so this day's move reflects our rubric, not a change to the server

### 2026-09-27 (score 78, +1)

No change was recorded against any check on this day. Stability & Change Management went from 83 to 87. That category is still filling its 30-day observation window: 25 days of observed history at the previous scan, 26 at this one. The score rises as the window fills, whether or not the server changes.

### 2026-09-25 (score 77, +3)

- [functional] We updated how we score, so this day's move reflects our rubric, not a change to the server

### 2026-09-24 (score 74, +1)

No change was recorded against any check on this day. Stability & Change Management went from 73 to 77. That category is still filling its 30-day observation window: 22 days of observed history at the previous scan, 23 at this one. The score rises as the window fills, whether or not the server changes.

### 2026-09-22 (score 73, +1)

No change was recorded against any check on this day. Stability & Change Management went from 67 to 70. That category is still filling its 30-day observation window: 20 days of observed history at the previous scan, 21 at this one. The score rises as the window fills, whether or not the server changes.

### 2026-09-20 (score 72, +1)

No change was recorded against any check on this day. Stability & Change Management went from 60 to 63. That category is still filling its 30-day observation window: 18 days of observed history at the previous scan, 19 at this one. The score rises as the window fills, whether or not the server changes.

### 2026-09-18 (score 71, +1)

No change was recorded against any check on this day. Stability & Change Management went from 53 to 57. That category is still filling its 30-day observation window: 16 days of observed history at the previous scan, 17 at this one. The score rises as the window fills, whether or not the server changes.

### 2026-09-15 (score 70, +1)

No change was recorded against any check on this day. Stability & Change Management went from 43 to 47. That category is still filling its 30-day observation window: 13 days of observed history at the previous scan, 14 at this one. The score rises as the window fills, whether or not the server changes.

## MCP tools (33)

### `list_windows` (~179 tokens)

Every open window with id, wm_class, title, geometry, pid and which one has focus. Start here: ids from this list are what type_text and press_keys target (ids change when a dialog is recreated -- a wm_class or title fragment does not). HOW TO DRIVE THIS DESKTOP, because the round trip is the expensive part and the actions are milliseconds: (1) ui_find then ui_press where the app has an accessibility tree -- it cannot miss; (2) find_text for Chrome, Electron and Qt, which expose almost nothing; (3) do_steps when you already know the next few actions, instead of one call each; (4) let the acting tool show you the result rather than following it with a screenshot -- they all do now. A screenshot of the whole screen is the last resort, not the first move.

### `screenshot` (~419 tokens)

Look at the screen, one window, or one rectangle. The image comes back in this reply -- there is nothing to Read afterwards. CROP, DO NOT SHRINK: `window` costs about 1300 tokens and a `region` strip about 160, against 1843 for the whole desktop, and all three stay legible, while `scale` below 1 makes small text unreadable for a saving a crop would have made anyway. Passing `window` AND `region` means a rectangle measured inside that window. `annotate` draws grid lines and window boxes labelled in SCREEN coordinates, so the number to pass to pointer_click can be read off the picture instead of estimated. Before reaching for this at all: ui_find and find_text answer "where is X" without an image, and every acting tool already shows you the result.

Input parameters:

- `annotate`: true for grid + window boxes, or an object: {grid: true|<spacing px>, windows: bool, widgets: bool, limit: int}.
- `app` (string): AT-SPI application name for annotate.widgets, if the focused window is not the one to map.
- `include_cursor` (boolean)
- `inline` (boolean): Set false to only write the PNG and get its path back, for a capture nobody needs to look at now.
- `path` (string): Where to write the PNG, e.g. /tmp/shot.png
- `region` (object): Capture just this rectangle, in screen pixels.
- `scale` (number): Resize the image you are shown. Below 1 it shrinks -- measured on this 1920x1080 screen, 0.5 makes UI text hard to read and OCR fails on it entirely, so crop instead. Above 1 it enlarges, which is wh…
- `window`: Capture just this window (id, wm_class or title fragment). Captures what is ON SCREEN there, so anything in front of it is included.

### `zoom` (~153 tokens)

Look closer at a small area at FULL resolution -- never scaled, unlike screenshot, which fits everything to the model's 1568px ceiling. For a tiny glyph, a hairline border, an icon. Refuses more than half the desktop: zoom exists to spend tokens on FEW pixels.

Input parameters:

- `pad` (integer): Extra pixels of context on every side.
- `path` (string): Where to keep the PNG; defaults to the shot cache
- `region`: Rectangle in screen pixels, [x, y, width, height] or {x, y, width, height}. With window=, measured inside that window.
- `window`: Zoom into this window (id, wm_class or title fragment).

### `pointer_move` (~286 tokens)

Move the pointer to an absolute screen position. Exact: this goes to the compositor (org.gnome.Mutter.RemoteDesktop), not through ydotool, so there is no acceleration curve and no closed loop needed. Coordinates are the same ones list_windows and screen_map report.

Input parameters:

- `expect_window`: Refuse the move if this window is not the one at that point.
- `look`: What to show you afterwards. Default "auto": wait for the screen to stop changing, measure how much this action changed, and attach a picture of the affected window only if something did change -- so…
- `look_at`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `settle_max_s` (number): How long to wait for the screen to stop changing before looking. Raise it for an app that animates slowly; set it to 0 to capture immediately.
- `x` (number, required)
- `y` (number, required)

### `pointer_click` (~558 tokens)

Click at an absolute screen position. Reports whether it LANDED on anything -- the screen is compared before and after, so a click into dead space says so instead of looking exactly like one that worked -- and shows you the result without a separate screenshot. Also reports whether the keyboard moved as a result. PASS expect_window: the click is refused if something else is under that point, which is the difference between a missed click and a click in someone else's window. Needs no consent dialog, unlike xdotool.

Input parameters:

- `button` (string)
- `count` (integer): 2 for a double click.
- `expect_window`: The window this click is aimed at (id, wm_class or title fragment). Nothing is clicked if it is not the window at that point.
- `hover_first` (boolean): Approach the point and settle before clicking, so a toolkit that only arms a button on hover gets its motion event. Chromium/CEF/Electron buttons (Creative Cloud, Spotify, 'desktop web' apps) commonl…
- `look`: What to show you afterwards. Default "auto": wait for the screen to stop changing, measure how much this action changed, and attach a picture of the affected window only if something did change -- so…
- `look_at`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `on_occluded` (string): What to do when expect_window is not the window at that point. Default refuses and names the blocker with its id and geometry. "click_topmost" clicks whatever is in front instead, in this same call,…
- `ref` (integer): A widget number from the last screen_map -- the click lands at that widget's CURRENT position after its identity is re-verified, no coordinates needed. Give ref OR x/y, never both. Refs die at the ne…
- `settle_max_s` (number): How long to wait for the screen to stop changing before looking. Raise it for an app that animates slowly; set it to 0 to capture immediately.
- `x` (number)
- `y` (number)

### `pointer_drag` (~353 tokens)

Press at one point, travel, release at another. The travel is real intermediate motion, because a press-and-teleport is not a drag to most toolkits.

Input parameters:

- `button` (string)
- `dwell_ms` (integer): Hover at the destination this long before releasing. Default 0 is the timing measured at 5/5 on a real drop target; a slower variant scored 4/5, so this is NOT a better default. Try ~400 only after a…
- `expect_window`
- `from_x` (number, required)
- `from_y` (number, required)
- `look`: What to show you afterwards. Default "auto": wait for the screen to stop changing, measure how much this action changed, and attach a picture of the affected window only if something did change -- so…
- `look_at`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `settle_max_s` (number): How long to wait for the screen to stop changing before looking. Raise it for an app that animates slowly; set it to 0 to capture immediately.
- `steps` (integer)
- `to_x` (number, required)
- `to_y` (number, required)

### `pointer_scroll` (~249 tokens)

Wheel clicks at a point. dy positive scrolls down, dx positive scrolls right.

Input parameters:

- `dx` (integer)
- `dy` (integer)
- `expect_window`
- `look`: What to show you afterwards. Default "auto": wait for the screen to stop changing, measure how much this action changed, and attach a picture of the affected window only if something did change -- so…
- `look_at`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `settle_max_s` (number): How long to wait for the screen to stop changing before looking. Raise it for an app that animates slowly; set it to 0 to capture immediately.
- `x` (number, required)
- `y` (number, required)

### `pointer_position` (~54 tokens)

Where the pointer is. Answers from the compositor when the extension supports it, otherwise from the last position this server set and says which. Never guesses from X, whose answer is stale whenever the pointer is over a Wayland surface.

### `window_at` (~73 tokens)

What a click at this point would hit. Use it before clicking somewhere you inferred from a screenshot. Reports both the compositor's own pick (which respects input shapes, so a click-through overlay is seen through) and every window whose rectangle covers the point.

Input parameters:

- `x` (number, required)
- `y` (number, required)

### `screen_map` (~156 tokens)

Everything on screen with the coordinates to reach it: the desktop rectangle, every window top of the stack first with its centre point, where the pointer is, and every pressable widget of the focused application with the exact pixel to click it at. This is the one call that turns 'click the Save button' into a number without looking at an image. Every widget also carries a ref: N -- pass it straight to ui_press(ref) or pointer_click(ref); refs are valid until the next screen_map call, and the result's refs_generation says which call issued them.

Input parameters:

- `app` (string): AT-SPI application name, if not the focused window's.
- `limit` (integer)
- `widgets` (boolean)

### `wait_for` (~231 tokens)

Wait until the desktop reaches a state, instead of sleeping a guessed number of seconds. Conditions: window_exists, window_gone, window_focused, focus_changes; text_appears (OCR polls a window for a string -- a reply arriving, a build finishing); widget_exists (an AT-SPI widget matching text/role shows up in app); clipboard_changed (a copy landed); elapsed (just wait N seconds -- for a long install with nothing to poll, and the honest alternative to watching for a string you know will never appear). Returns as soon as it is true, or reports honestly that it timed out. A timeout over 300s is clamped, not refused.

Input parameters:

- `app` (string): For widget_exists: the AT-SPI application name
- `condition` (string, required)
- `role` (string): For widget_exists: the widget role, e.g. 'push button'
- `target`
- `text` (string): For text_appears: the string to watch for. For widget_exists: the widget name to match.
- `timeout` (number)

### `assert_state` (~177 tokens)

Prove the desktop is in a state, with evidence -- the honest way to END a task. Each assertion comes back passed/failed with what was actually observed; a false assertion is a result, not an error. Give any of: window_exists, window_focused, text_present, widget_exists, clipboard_contains.

Input parameters:

- `clipboard_contains` (string): The clipboard text must contain this string.
- `text_present` (object): {"text": ..., "window": ...}: the string must be visible (OCR) in that window.
- `widget_exists` (object): {"app": ..., "text" and/or "role"}: a matching AT-SPI widget must exist.
- `window_exists`: A window id, wm_class or title fragment that must exist.
- `window_focused`: A window that must exist AND hold focus.

### `find_text` (~260 tokens)

Where a visible piece of text is on screen, in coordinates you can click. Reads the pixels with OCR, so it works in Chrome, Electron and Qt apps, which expose almost nothing to ui_find. Try ui_find FIRST -- pressing a real widget cannot miss -- and come here when it returns nothing. About 1.5s for a window; cheaper and more exact than taking a picture and estimating. Blind to icon-only buttons: there is no text in them to read.

Input parameters:

- `exact` (boolean): Whole word, case sensitive.
- `limit` (integer)
- `min_confidence` (integer)
- `psm` (integer): tesseract page segmentation mode. Defaults to 6 inside a window and 11 for the whole screen, which is what measured best for each.
- `region`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `text` (string, required): The visible string to find. A phrase is matched across consecutive words on one line.
- `window`: Search inside this window only (id, wm_class or title fragment). Much faster and far fewer false matches than the whole screen.

### `region_changed` (~135 tokens)

Wait until a window or rectangle CHANGES, then show it. For anything wait_for cannot express: a reply arriving, a spinner finishing, a download completing. Polls pixels here instead of making you take blind screenshots and look at each one, and returns as soon as it changes.

Input parameters:

- `look` (boolean): Attach the picture once it changes.
- `poll_seconds` (number)
- `region`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `timeout` (number)
- `window`

### `do_steps` (~516 tokens)

Run a short sequence of actions in ONE call and look once at the end. Every separate tool call costs a model round trip of several seconds while the action itself takes milliseconds, so a known sequence -- activate, click, type, press Return, see the result -- belongs here rather than in four calls. Steps run with their own `look` off; the picture is taken after the last one, or at the step that failed. Use single tools when the next action depends on what the last one revealed. The sequence is validated up front -- a call that cannot finish never starts.

Input parameters:

- `look`: What to show you afterwards. Default "auto": wait for the screen to stop changing, measure how much this action changed, and attach a picture of the affected window only if something did change -- so…
- `look_at`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `settle_max_s` (number): How long to wait for the screen to stop changing before looking. Raise it for an app that animates slowly; set it to 0 to capture immediately.
- `steps` (array, required): Ordered actions. Each is {do: verb, ...the same arguments that tool takes}. Verbs: activate(target), click(x,y), move(x,y), drag(from_x,from_y,to_x,to_y), scroll(x,y,dy), type(target,text), key(targe…
- `stop_on_error` (boolean): Stop at the first failing step. Leave this true unless the steps are genuinely independent.

### `screencast` (~204 tokens)

Record the screen, or one window, to an h264 mp4. Use this instead of screenshot whenever the thing being judged MOVES -- an animation, a transition, a scroll, a stutter, a hover state. Stills cannot show motion and bursting them tops out near 5 fps. Goes under the xdg portal straight to org.gnome.Mutter.ScreenCast, so there is no share-your-screen consent dialog, and encodes on the iGPU. Read the result back by pulling frames out with ffmpeg.

Input parameters:

- `fps` (integer)
- `include_cursor` (boolean)
- `path` (string, required): Where to write the mp4, e.g. /tmp/cast.mp4
- `seconds` (number)
- `target`: Window id from list_windows, or a wm_class / title fragment. The window is activated and focus is CONFIRMED before any key is sent; if focus does not land, nothing is typed.

### `frames` (~268 tokens)

Turn a video into ONE image you can actually look at: N frames, evenly spaced, stamped with frame number and timestamp, tiled into a contact sheet. This is the other half of screencast -- a model cannot decode an mp4, so a recording is useless until it becomes stills. Also measures per-frame change and reports duplicate frames, which detects a source repainting slower than the capture rate and catches stutter and frozen output that eyeballing misses. Use from_frame/to_frame to zoom into a fraction of a second once the overview shows where the interesting moment is. Works on any video, not just screencast output.

Input parameters:

- `cols` (integer)
- `compare` (string): A second video. Its sheet is stacked underneath the first in ONE image, which is what a before/after needs -- two separate sheets are never on screen together to be compared.
- `from_frame` (integer): Start of a dense slice, in frames. Omit to span the whole clip.
- `outdir` (string): Where to write the sheet; defaults to <video>-frames next to the video
- `path` (string, required): The video to read, e.g. /tmp/cast.mp4
- `rows` (integer)
- `to_frame` (integer)

### `activate_window` (~263 tokens)

Focus and raise a window, then confirm focus actually landed there. Returns an error rather than a false success.

Input parameters:

- `look`: What to show you afterwards. Default "auto": wait for the screen to stop changing, measure how much this action changed, and attach a picture of the affected window only if something did change -- so…
- `look_at`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `settle_max_s` (number): How long to wait for the screen to stop changing before looking. Raise it for an app that animates slowly; set it to 0 to capture immediately.
- `target` (required): Window id from list_windows, or a wm_class / title fragment. The window is activated and focus is CONFIRMED before any key is sent; if focus does not land, nothing is typed.

### `window_manage` (~163 tokens)

Move, resize, close, (un)minimize, (un)maximize, re-workspace or pin a window -- through the compositor, where these are ordinary calls. The result reports the window as it IS afterwards (new geometry, or gone), not just that the call was sent. close that leaves the window standing names the usual reason: an unsaved-changes dialog.

Input parameters:

- `above` (boolean): For action: above -- pin or unpin.
- `action` (string, required)
- `height` (integer)
- `index` (integer): Workspace index, for action: workspace
- `target` (required): Window id, wm_class or title fragment.
- `width` (integer)
- `x` (integer)
- `y` (integer)

### `launch_app` (~176 tokens)

Start an application and confirm it actually arrived: the result carries the NEW window's dict (or, while the screen is locked, the new AT-SPI app) and names which mechanism confirmed. Every real task starts with an app that is not running yet; this is that step, inside the protocol instead of a shell command.

Input parameters:

- `command` (array): Argv list, not a shell string. Exactly one of desktop_id/command.
- `desktop_id` (string): Desktop id for `gio launch`, with or without '.desktop', e.g. 'org.gnome.TextEditor'
- `file` (string): Optional file for the app to open
- `timeout` (number)
- `wait_window` (boolean): Confirm arrival by a NEW window id (or a new AT-SPI app when the extension is down).

### `ui_apps` (~28 tokens)

Applications currently on the AT-SPI bus. These names are what ui_tree and ui_find take.

### `ui_tree` (~60 tokens)

Accessibility tree for one application: roles, names, screen bounds, and which nodes are actionable. Prefer ui_find unless you genuinely need the shape of the whole window.

Input parameters:

- `app` (string, required): Application name from ui_apps
- `depth` (integer)

### `ui_find` (~176 tokens)

Find widgets by visible text. THE way to locate something to act on: pressing a real widget through AT-SPI cannot miss and does not care where the window moved to. Paths returned here are valid only while the tree is unchanged -- find, then act.

Input parameters:

- `actionable_only` (boolean): Only widgets that expose an action
- `app` (string): Restrict to one application (much faster)
- `depth` (integer): GTK4 nests deeply -- the default is 30 because a text view can sit at depth 23
- `role` (string): Require an exact AT-SPI role, e.g. push_button
- `text` (string): Substring to look for in widget names. Optional if role or actionable_only is given -- which is how you list the icon-only buttons that have no name to search.

### `ui_read_text` (~101 tokens)

Read the content of a text widget straight out of the accessibility tree. This is how you VERIFY that something landed, instead of trusting that a keystroke arrived.

Input parameters:

- `app` (string): Application name; its editable text widget is located automatically
- `path` (string): Or an exact index path. Both tools return the resolved path -- pass it back to address the SAME document across write and read; without it, both prefer the focused text widget.

### `ui_set_text` (~345 tokens)

PREFERRED way to enter text. Writes through AT-SPI EditableText, which needs no focus and no ydotool: it works on an unfocused window and even while the screen is locked, and it reads the widget back to prove the text landed. Use type_text only when a widget is not AT-SPI-editable.

Input parameters:

- `app` (string): Application name; its editable text widget is located automatically
- `look`: What to show you afterwards. Default "auto": wait for the screen to stop changing, measure how much this action changed, and attach a picture of the affected window only if something did change -- so…
- `look_at`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `path` (string): Or an exact index path. Both tools return the resolved path -- pass it back to address the SAME document across write and read; without it, both prefer the focused text widget.
- `replace` (boolean): Clear existing content first
- `settle_max_s` (number): How long to wait for the screen to stop changing before looking. Raise it for an app that animates slowly; set it to 0 to capture immediately.
- `text` (string, required): Text to write

### `ui_press` (~346 tokens)

Invoke a widget's own action through AT-SPI -- the preferred way to act on this desktop. Requires expect_name or expect_role, and refuses if the path no longer points at that widget, so a shifted tree cannot make you press the wrong thing.

Input parameters:

- `action_index` (integer)
- `expect_name` (string): Name the widget should still have (substring)
- `expect_role` (string): Role the widget should still have
- `look`: What to show you afterwards. Default "auto": wait for the screen to stop changing, measure how much this action changed, and attach a picture of the affected window only if something did change -- so…
- `look_at`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `path` (string): Index path from ui_find, e.g. "gedit/0/3/1"
- `ref` (integer): A widget number from the last screen_map; path, expect_name and expect_role are filled from it. Give ref OR path, never both.
- `settle_max_s` (number): How long to wait for the screen to stop changing before looking. Raise it for an app that animates slowly; set it to 0 to capture immediately.

### `type_text` (~406 tokens)

Type into a named window. Focus is confirmed first, nothing is typed if it cannot be confirmed, and the widget is read back afterwards to check the right characters arrived. Characters go to the compositor as keysyms, so the keyboard layout cannot transpose them -- the German-QWERTZ hazard that made ydotool type z for y does not apply to this path. ui_set_text is still better where it works: it hands text to the widget and needs no focus at all.

Input parameters:

- `key_delay_ms` (integer)
- `look`: What to show you afterwards. Default "auto": wait for the screen to stop changing, measure how much this action changed, and attach a picture of the affected window only if something did change -- so…
- `look_at`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `settle_max_s` (number): How long to wait for the screen to stop changing before looking. Raise it for an app that animates slowly; set it to 0 to capture immediately.
- `target` (required): Window id from list_windows, or a wm_class / title fragment. The window is activated and focus is CONFIRMED before any key is sent; if focus does not land, nothing is typed.
- `text` (string, required): Literal text to type
- `verify_app` (string): AT-SPI application name to read back for verification; auto-detected from the window if omitted
- `via` (string): auto prefers compositor keysyms and falls back to ydotool.

### `clipboard_write` (~149 tokens)

Put text or a file's bytes on the clipboard, and PROVE it landed by reading it back. Goes through the gnome-shell extension so the compositor sets the clipboard itself -- mutter has a measured bug (S-018) where an external client's offer can serve wrong bytes to text requests, so wl-copy is only the fallback and says so when used.

Input parameters:

- `mimetype` (string): The type those bytes are offered as, e.g. image/png. Required with path.
- `path` (string): File whose bytes go on the clipboard, e.g. a PNG
- `text` (string): Text to place on the clipboard. Give this OR path+mimetype, not both.

### `clipboard_read` (~75 tokens)

What is on the clipboard. Text by default; types:true lists the offered mimetypes instead. An empty clipboard is a clean result, not an error, and a clipboard owner that never serves its offer is reported after a short deadline instead of hanging.

Input parameters:

- `types` (boolean): List offered mimetypes instead of reading text.

### `press_keys` (~317 tokens)

Send a key combination to a named window, e.g. ctrl+s. Chain several with do_steps rather than one call each. Focus is confirmed first. Ctrl+Alt+F1-F12 is refused: it switches virtual terminal and looks exactly like a frozen machine.

Input parameters:

- `combo` (string, required): e.g. 'ctrl+shift+t'
- `look`: What to show you afterwards. Default "auto": wait for the screen to stop changing, measure how much this action changed, and attach a picture of the affected window only if something did change -- so…
- `look_at`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `settle_max_s` (number): How long to wait for the screen to stop changing before looking. Raise it for an app that animates slowly; set it to 0 to capture immediately.
- `target` (required): Window id from list_windows, or a wm_class / title fragment. The window is activated and focus is CONFIRMED before any key is sent; if focus does not land, nothing is typed.
- `via` (string)

### `hold_key` (~331 tokens)

Hold ONE key down for a duration, then release it -- a real press and a separate release, not a tap. For shift-selection, held-key scrolling and games. Blocks for the whole duration; press_keys is the tool for combinations.

Input parameters:

- `key` (string, required): One key, e.g. 'shift', 'down', 'w'. No combos.
- `look`: What to show you afterwards. Default "auto": wait for the screen to stop changing, measure how much this action changed, and attach a picture of the affected window only if something did change -- so…
- `look_at`: Rectangle for look:"region", in screen pixels. Object form {x, y, width, height} or array form [x, y, width, height].
- `seconds` (number, required): How long to hold. The call blocks for this long.
- `settle_max_s` (number): How long to wait for the screen to stop changing before looking. Raise it for an app that animates slowly; set it to 0 to capture immediately.
- `target`: Window id from list_windows, or a wm_class / title fragment. The window is activated and focus is CONFIRMED before any key is sent; if focus does not land, nothing is typed.

### `desktop_health` (~74 tokens)

Whether each mechanism is usable right now, and what each one will actually do: extension state and which methods the RUNNING shell has (an edited extension does not load until the next login), absolute pointer control, window and AT-SPI counts, keyboard layout, and the XTEST trap. Call this first when something behaves oddly.

### `journal` (~108 tokens)

Read back the trail of acted tool calls -- every state-changing call is journaled with its arguments, outcome, hit/miss verdict and screenshot hash. Use it to reconstruct what already happened after context loss, or to review an unattended run. Reading tools are not in it.

Input parameters:

- `session` (string): current: only this server process's actions. all: every session in the 14-day retention window.
- `tail` (integer): How many of the most recent entries to return, oldest first.

## Diagnostics

Captured diagnostic sections: Provenance, Install scripts, Dependencies. The full working is on the page: https://verifymcp.io/servers/tristanmuzzu-deskwright/deskwright#diagnostics

## Score history

- 2026-09-28: 89
- 2026-09-27: 78
- 2026-09-26: 77
- 2026-09-25: 77
- 2026-09-24: 74
- 2026-09-23: 73
- 2026-09-22: 73
- 2026-09-21: 72
- 2026-09-20: 72
- 2026-09-19: 71
- 2026-09-18: 71
- 2026-09-17: 70
- 2026-09-16: 70
- 2026-09-15: 70
- 2026-09-14: 69
- 2026-09-13: 69
- 2026-09-12: 68
- 2026-09-11: 68
- 2026-09-10: 67
- 2026-09-09: 67
- 2026-09-08: 63
- 2026-09-07: 63
- 2026-09-06: 63
- 2026-09-05: 63
- 2026-09-04: 63
- 2026-09-03: 63
- 2026-09-02: 63
- 2026-09-01: 63

## Common questions

### What is the Deskwright MCP server?

Deskwright is an MCP server listed in the public MCP registry as io.github.tristanmuzzu/deskwright. Agents use real GNOME/Wayland apps: on your screen, or an invisible second desktop you never see. This page covers its PyPI package (deskwright).

### Is the Deskwright MCP server safe to use?

Deskwright scores 89 out of 100 on VerifyMCP. We found no known CVEs affecting it as of 28 September 2026. Its build provenance is signed and verified. That is a record of what we were able to check automatically, not an endorsement. The category breakdown on this page shows every signal behind the number, including the ones we could not confirm.

### What tools does the Deskwright MCP server expose?

Deskwright exposes 33 tools: list_windows, screenshot, zoom, pointer_move, pointer_click, and 28 more. Their descriptions and schemas cost roughly 7,389 tokens of context every time the server is loaded.

### Is the Deskwright MCP server still maintained?

Deskwright is still listed as active in the MCP registry. We last reached this channel on 28 September 2026. Those dates come from our own scans of the registry and the channel itself, not from anything the publisher announced.

### What licence is the Deskwright MCP server under?

Deskwright declares the Apache-2.0 licence, which is OSI-approved. That covers the source only, and says nothing about the cost of any service it calls.

## Links

- PyPI project: https://pypi.org/project/deskwright/
- Socket report: https://socket.dev/pypi/package/deskwright
- Repository: https://github.com/tristanmuzzu/deskwright
- Changelog RSS feed: https://verifymcp.io/servers/tristanmuzzu-deskwright/deskwright.xml
- Changelog JSON feed: https://verifymcp.io/servers/tristanmuzzu-deskwright/deskwright.json
- HTML version of this page: https://verifymcp.io/servers/tristanmuzzu-deskwright/deskwright
