Case study
02 A desktop app for local AI coding agents
Alfredo
Give AI coding agents a clear task and inspect their changes before a pull request.
- What I built
- A local desktop workstation that runs agents in isolated working copies, collects changes and tests, and supports review and repair.
- My contribution
- My work connects task boundaries, agent execution, and evidence review. AI-assisted implementation is part of the process; responsibility for scope and acceptance stays explicit.
- Important limit
- The evidence does not establish a public release, broad operating-system compatibility, production performance, or completed human accessibility acceptance.
Workflow recording
The sample task adds a README Verification section with three checklist items: clear scope, no personal data and no dependencies. The local agent changes only README.md. Its actual diff is inspected and the native review is approved.
Unreleased migrated build with an isolated schema and prompt-framing repair; recorded 8 September 2026. The task review succeeds. Automatic workspace cleanup is separately blocked by this host; the preserved snapshot remains intact.
For developers using AI coding agents, knowing what changed matters as much as getting a result.
- What works
- Recorded integration checks cover task dispatch, review, controlled crash recovery, restart continuity, and an installed Linux candidate.
- Key decision
- Permit automatic actions within recorded authority; escalate decisions that need more permission. Accepted work is PR-ready.
- Next milestone
- Complete public-registry reinstall, visible-desktop and assistive-technology review, and repeated performance measurements.
- Snapshot
- Public repository evidence / main at af74a85
- Evidence date
- 2026-09-03
Intent
Why this work exists. The problem worth solving and the human judgment that frames it.
Make an agent's work inspectable
A useful-looking change can hide which files an agent was allowed to touch, which commands it ran, and how the result was accepted. Alfredo brings the task, permissions, changes, tests, and review outcome together in one desktop app. Its Python backend checks and records authorized actions.
Where a person makes the decision
People select the repository and Mission and handle permission requests or escalated reviews. Matching tasks may be approved automatically within existing authority. People can review evidence, and the architecture also gives a model reviewer a role in repair and PR-readiness decisions. The backend checks those decisions; incomplete evidence cannot be accepted. Merge approval remains separate, and this portfolio's publication approval belongs to its owner.
- Problem
- When an AI agent changes code, you need to know what it was allowed to change, what it actually changed, which tests ran, and how the result was accepted.
- Constraints
- Keep task state, permissions, execution, evidence, and review decisions in the Python Orchestrator, the backend that checks and records authorized actions. Bind each delegated slice to an exact Mission, goal, acceptance criteria, allowed paths, command policy, and eligible Local Agent before launch. Run agents in isolated working copies with limits on files, commands, and resources. Incomplete evidence can be inspected, but acceptance requires a complete Evidence Package. Allow automatic actions only within the recorded permission boundary; gated or uncertain decisions return to a person. Accepted work is PR-ready and does not authorize a merge.
Build
How the work was shaped. The workflow, tailoring, present capability, and next useful step.
From a chosen repository to a reviewed result
Select the repository, called the Coding Workspace, and start or resume a Mission, the recorded body of work. Alfredo checks the task and agent against existing permissions before running it in an isolated working copy. Changes and tests form an Evidence Package for inspection. Review can accept a complete result, request repair, or escalate to a person. Accepted work is ready for a pull request handoff.
From review rules to the connected app
The sources record issue-sized implementation work. The June slice adds backend review rules, typed client and desktop interfaces, and rendered tests. The August work connects the React clients to the real Python backend and separately verifies the installed Linux candidate, recovery, restart, and browser layouts. This connects focused review-rule tests to checks of the assembled app and its installed package.
Local agents, recorded permissions, visible results
The reviewed architecture describes a React/Tauri desktop app with a Python Orchestrator, the backend responsible for task state and permission checks. It supports tracked and ad hoc tasks, isolated worktrees, evidence inspection, and review and repair. The August integration report exercises dispatch, acceptance, controlled crash recovery, retained session records, and restoration of the work history after restart.
Finish release and human usability checks
The August report leaves public package publication and a fresh public-registry installation open. It also calls for people to test the installed desktop with a keyboard and assistive technology, plus repeated measurements of the installed app before making performance claims. These are the report's remaining checks at the reviewed snapshot.
Project workflow
- Choose a repository and task — Select the Git repository, called the Coding Workspace, then start or resume a Mission: the recorded body of work.
- Check the permitted scope — Verify the goal, success criteria, allowed files, commands, and eligible agent. Matching requests can proceed under existing authority; other requests need a decision.
- Run an agent in isolation — Queue a local coding agent in an isolated worktree, a separate working copy of the repository.
- Collect changes and test results — Build an Evidence Package containing changed files, a diff summary, commands, test results, risks, proposed context updates, and inspectable artifacts.
- Review and repair — Inspect complete or incomplete evidence. Review decisions can accept the result, request repair, or escalate to a person; incomplete packages cannot be accepted.
- Prepare a pull request handoff — Keep accepted evidence and prepare PR-ready instructions. Acceptance does not authorize merging the changes.
Development workflow
- Build in issue-sized pieces — The architecture records a development backlog organized into Issue Slices, each representing a piece of the product.
- Implement and test review rules — The June report records Python review decisions, typed client and desktop contracts, and tests for incomplete evidence and stale decisions.
- Test the connected workstation — The August report drives the React clients against the real Python backend through repository selection, task dispatch, review, recovery, and restart.
- Verify the installed Linux candidate — A separate gate builds and installs the packaged desktop through a local test registry, then checks the launcher, interface, backend, and artifact integrity.
- Record results and remaining checks — The reports list passing checks alongside unfinished release, human accessibility, and performance work.
Proof
What the evidence supports. The demonstration worth inspecting and the boundary it cannot cross.
What the reports show
The June report records tests of evidence inspection, acceptance, repair, and escalation to a person. The August report drives real backend actions through test clients and separately checks an installed Linux candidate and responsive browser layouts. These are documented test results, not a human production-usage study. The recording shows a sample README task in the unreleased migrated workstation: task entry, local execution, diff inspection and review approval. Processing waits are shortened. An isolated schema and prompt-framing repair was used; automatic workspace cleanup remains separately blocked on the capture host. The synthetic workflow evidence is an illustration, not another captured run.
The limits of the recorded evidence
The sources support the local workstation and a tested Linux release candidate. They do not establish public package availability, broad operating-system compatibility, completed human desktop or accessibility testing, or production performance. Accepted changes reach PR-ready status without merge approval. The portfolio owner separately decides whether to publish this account.
Synthetic workflow evidence
- Before / Boundary approved — Task contract
Allowed: guide.md · Worktree: Isolated · Bounded / Ready to delegate
- Decision / Boundary check fails — Evidence Package
Expected: guide.md · Unexpected: settings.json · Invalid / Invalid result
- After / Human-led review accepts repair — Review outcome
Boundary: Restored · Authority: Human · Accepted / PR-ready, not merged
Named human decisions
- Set the scope and resolve permission requests
A person selects the repository and Mission. Alfredo checks each delegation against recorded authority; requests requiring additional permission or a gated worker need an explicit decision.
Why: Automatic approval is possible within an existing boundary. The routing model cannot expand that boundary by proposing a task.
- Handle decisions that need human review
A person can inspect the evidence and choose acceptance or repair, and handles outcomes escalated for human review. The architecture also supports a model reviewer whose recorded outcomes can drive repair and PR readiness.
Why: Raw test output or model prose alone cannot authorize acceptance; the Orchestrator checks the review decision and evidence requirements.
- Keep delivery separately authorized
Alfredo prepares pull request instructions without granting merge approval. Its documented package publication path requires separate authorization. Approval to publish this portfolio narrative is the portfolio owner's policy.
Why: A recorded review outcome does not establish that a package was released or that the owner approved a public portfolio claim.
Validation provenance / status
- Modernized workstation verification / 2026-08-30
Evidence date: 2026-08-30 · Status: recorded · Environment: Real Python CLI and backend, React/Tauri workstation, installed Linux release candidate, and Chromium desktop-to-mobile viewports · Evidence type: Public-seam integration, recovery, release-candidate, and responsive-browser evidence
Result: The recorded journey selects an exact workspace and Mission, confirms a shared-understanding gate, dispatches governed tracked and ad hoc work, reviews evidence, recovers a controlled crash, retires accepted sessions, and restores the same chronology after restart.
Provenance: Public Issue #75 modernized-workstation verification report at the reviewed Alfredo main snapshot.
Material limitations: The run did not publish packages, perform registry-only reinstall, complete human real-display or assistive-technology acceptance, or establish production performance.
- Review Workspace Evidence Package / 2026-06-25
Evidence date: 2026-06-25 · Status: recorded · Environment: Python Orchestrator, React and TypeScript clients, and Rust/Tauri bridge test surfaces · Evidence type: Evidence contract, integration, and rendered review-action tests
Result: The focused record verifies complete and incomplete Evidence Packages plus acknowledged accept, repair, and human-escalation decisions, including stale-revision protection and disabled incomplete acceptance.
Provenance: Public Review Workspace Evidence Package completion report at the reviewed Alfredo main snapshot.
Material limitations: This focused historical record supports the review seam; the later integrated workstation report is the current overall capability snapshot.
- Alfredo: Reviewed repository snapshot — Source code, documentation, tests, and project history for the local coding-agent workstation.
- Alfredo: Architecture — Authoritative Mission, isolation, execution, Evidence Package, review, repair, and delivery boundaries.
- Alfredo: Evidence Package review — Focused implementation and test evidence for accept, repair, and human-escalation review decisions.
- Alfredo: Integrated verification — Current public-seam workstation, recovery, installed-candidate, responsive-browser, and open-acceptance evidence.
Bounded proof: The reviewed evidence supports a local workstation and verified Linux release candidate; it does not prove public registry publication, broader operating-system compatibility, human real-display or assistive-technology acceptance, production performance, or automatic merge or deployment.