Inside PixelSteer

A little context goes a long way.

PixelSteer adds a visual feedback layer to your running frontend. You point at the interface; your coding agent changes the actual source.

Already in an agent session?

Give it the skill. Keep the conversation.

Install the skill and invoke it in the session you already have open:

npx skills add PixelSteer/pixelsteer-skill

Then ask your coding agent from your frontend project:

Use the pixelsteer skill to start the visual feedback session for my frontend. Keep handling browser feedback until I ask you to stop.

The skill handles the project-specific setup. Use this route inside an existing agent session; the launch flags are for starting a new session from your terminal.

PixelSteer checks write access to its persistent license directory before starting. If a sandbox blocks access, the skill retries through the agent’s approval mechanism with access to the directory reported in the startup error. This lets you activate a license later in the same session.

Start from your terminal

Pick your copilot. Bring your frontend.

If you are in your app’s folder, use a command below to start steering your agent visually. PixelSteer opens the agent with the feedback workflow in its prompt. The agent starts your frontend and PixelSteer, then gives you a browser URL.

Run from your frontend project with PixelSteer and your chosen agent CLI installed and available on PATH. Sign in to the agent first; it uses your existing account and permission settings.

Codex

Your terminal, ready to steer.

pixelsteer --launch-agent --codex

The launch prompt includes the feedback workflow, so you can start without installing a skill. Ask the agent to stop when you are finished so it can close the servers it started. Abruptly closing the agent can leave those servers running.

Manual setup · npm

Your scripts. Your starting point.

Keep PixelSteer beside your frontend’s other development tools. A small npm script gives you one entry point for the proxy, agent launch, and task commands.

01 · Install it in your project

From the frontend directory, add PixelSteer as a development dependency. The npm package requires Node.js 18 or newer and installs the native executable.

npm install --save-dev pixelsteer

Add this entry to the existing scripts object in your package.json, keeping your other scripts:

"pixelsteer": "pixelsteer"

Add .codex/tasks/ to your frontend’s .gitignore so local feedback and session files stay out of commits.

02 · Launch your agent through npm

With your chosen agent CLI installed and signed in, run:

npm run pixelsteer -- --launch-agent --codex

Choose Claude Code or supply your own command instead:

npm run pixelsteer -- --launch-agent --claude
npm run pixelsteer -- --launch-agent --agent='/path/to/agent {prompt}'
The extra -- is the handoff.

It tells npm to pass the remaining arguments to PixelSteer. Keep --launch-agent together with exactly one agent selector. The agent starts the frontend and proxy, verifies the browser URL, and enters the feedback loop.

03 · Or start the proxy yourself

For an existing agent session or a workflow you manage yourself, start your frontend with its usual command in one terminal. If that command is npm run dev, for example:

npm run dev

In a second terminal, point PixelSteer at the URL printed by your dev server. This example uses port 5173:

npm run pixelsteer -- --target http://localhost:5173 --port 3100

Open the PixelSteer URL printed in the terminal, normally http://localhost:3100. Keep both terminals running. The proxy captures feedback; an agent still needs to claim tasks and edit the source. Use the skill in your existing session or give the agent the task commands below.

Connect an existing agent with the task commands

Have the agent run these from the same frontend root as the proxy. Add --project-root /absolute/frontend to task commands if it works from another directory.

npm run pixelsteer -- tasks next --wait

next --wait waits quietly, claims one task, and returns its JSON context. The agent uses the returned ID and selections to edit and validate the project. While working, it can report progress:

npm run pixelsteer -- tasks progress TASK_ID --message "Adjusting the layout"

After a successful edit, record the result and wait for the next task:

npm run pixelsteer -- tasks complete TASK_ID --result "Updated and checked the layout" --changed-file src/App.tsx --wait

Replace TASK_ID and the changed-file path with the real values. Repeat --changed-file for multiple files. If the task cannot be completed, report the reason:

npm run pixelsteer -- tasks fail TASK_ID --result "Required configuration is missing" --wait

Keep one waiter per agent session and only update tasks it claimed. With --wait, a completion or failure acknowledgement is followed by waiting for the next task; resume that same command session. Omit --wait when finishing the final task and ending the session.

Browse the full command reference with npm run pixelsteer -- tasks --help. If Auto execute is off in the overlay, click Execute now to send the plan to the task queue.

When you finish a manually started session, stop the agent’s task waiter and press Ctrl + C in each server terminal you started.

The workflow

From visual note to source edit.

Your normal dev server keeps running. PixelSteer sits in front of it and gives the browser a small selection interface.

01Run your app
02Open PixelSteer
03Select and describe
04Agent edits source
05Review via HMR
  1. Start the project with its usual development command and PixelSteer.
  2. Open the PixelSteer URL. Use Browse normally, then switch to element or rectangle selection when you want to leave feedback.
  3. Select one or more parts of the UI and describe the result you want.
  4. Your coding agent claims the task, inspects the project, and edits the relevant source files.
  5. Your existing dev server detects the edit. HMR refreshes the page so you can review and repeat.

Behind the scenes

A short path from browser to source.

flowchart TB
  accTitle: PixelSteer feedback workflow
  accDescr: PixelSteer forwards browser requests to the existing dev server and injects the selection overlay. Browser feedback enters the local task queue. The coding agent claims tasks, edits project source, and reports status through PixelSteer to the browser.
  browser[Browser]
  proxy[PixelSteer reverse proxy]
  server[Existing dev server]
  overlay[Selection overlay]
  queue[Local task queue]
  agent[Coding agent with PixelSteer skill]
  source[Project source]
  browser -->|Page requests| proxy
  proxy <-->|Requests and live reload| server
  proxy -->|Injects into HTML| overlay
  overlay -->|Visual feedback| queue
  queue -->|Claims task| agent
  agent -->|Edits| source
  source -->|Rebuilds| server
  agent -->|Reports status| queue
  queue -->|Status updates| proxy
  proxy -->|WebSocket updates| browser

Reverse proxy

PixelSteer forwards your app and modifies only HTML responses to add its browser client. Assets, cookies, query strings, streaming responses, HMR sockets, and other upgrades continue through the proxy.

Isolated overlay

The controls render inside a Shadow DOM so the app's styles do not leak into PixelSteer—and PixelSteer's styles do not change the app.

Local task queue

Feedback becomes a JSON task under .codex/tasks/. Tasks move through pending, working, completed, or failed states.

Agent skill

The installed skill watches for work, claims each task atomically, edits the project with the coding agent, and records the result and changed files.

Live status

A filesystem watcher relays task progress to the overlay over PixelSteer's WebSocket connection.

Existing toolchain

PixelSteer does not replace your framework, dev command, build system, or hot-reload loop. It runs beside them.

Context

What the agent receives.

Each task combines your written request with structured information captured from the selected interface:

  • page URL, viewport size, and scroll position;
  • element tag, ID, classes, attributes, and visible text;
  • DOM path, bounding rectangle, and parent context;
  • relevant computed styles such as layout, spacing, type, color, borders, flex, and grid; and
  • a bounded HTML snapshot of each selected element.

The agent combines this evidence with the project itself to find the implementation and make the requested change.

Source of truth

Your code stays canonical.

No parallel page model.

PixelSteer does not rebuild your interface in a proprietary canvas or save visual changes somewhere else. The coding agent edits the same source files you already own, review, and commit.

What persists

The code changes persist because they are ordinary project edits. PixelSteer's task files provide a durable local handoff and history for the agent workflow.

What stays local

The proxy, browser connection, task queue, and agent handoff run on your machine. Start an agent with the launch flags or bring an existing session with the skill. Your coding agent may send context to its model provider according to its own settings.