From prompts to inspectable work
The architecture and thesis behind MachinArc: bounded Machines, visible workflows, and durable human oversight.
The thesis
A prompt can initiate useful work, but a sustained workflow needs more than a response. It needs a defined objective, a record of attempted actions, a way to stop or approve those actions, and a clear account of how the result was produced.
MachinArc explores an operating layer around AI execution. Its central objects are configured Machines, individual runs, and Arcs that coordinate work. The design makes execution state available to operators instead of asking them to reconstruct it from a final paragraph.
The unit of work
A Machine combines an objective and success criteria with a model, tool set, permission configuration, memory mode, and execution limits. A task supplies the input for a run of that configuration. Separating configuration from a particular run makes the intended behavior easier to inspect and revise.
An Arc composes these units into a directed workflow. Its nodes include Machine execution, branching decisions, human review, delays, and final output. A graph snapshot anchors each run to the workflow that created it, while node records expose progress and failure paths.
System architecture
The application uses Next.js, server actions, PostgreSQL, and Drizzle. Server-side application sessions establish operator identity, and workspace-scoped queries enforce access boundaries. Application tables are protected from public Data API roles; the server connects through a dedicated database role with explicit grants and row-level policies.
A PostgreSQL job queue separates submission from execution. Workers claim jobs using database locking and leases. A hosted scheduler invokes a bounded worker endpoint, while a dedicated process supports work that needs a longer execution window.
- Control plane: workspace configuration, sessions, permissions, integrations, and operator decisions.
- Execution plane: provider calls, tool execution, workflow advancement, and bounded workers.
- Persistence plane: tasks, runs, graph snapshots, approvals, results, usage, and job receipts.
Durability is a product property
The runtime persists messages, proposed tool calls, tool execution records, approval state, and results. Recovery reconciles unanswered calls before advancing the model conversation and reuses a completed result where one exists.
Trigger creation has a separate durability boundary. Schedule advancement and fire-job enqueueing are transactional. A stable trigger job receipt records the start outcome, and the created run, its driver job, and schedule counters commit together. A reclaimed delivery can refer to the same outcome without creating another run.
These mechanisms support restart recovery, but they do not make arbitrary external side effects exactly once. When an interrupted action has an unknown outcome, execution stops for reconciliation rather than replaying a change that may already have occurred.
Economics and resource boundaries
Model turns, tool calls, runtime, and budgets bound the work a task can attempt. Usage is recorded separately from the cost calculation, which depends on the application's configured model pricing. Delegation adds child work to the parent execution story.
Workspace Credits and network interactions are prototype accounting concepts. They are not evidence of settled customer payments, an open provider marketplace, or validated commercial unit economics. Live provider costs, customer willingness to pay, and support requirements still need measurement.
Current state
MachinArc is a working prototype with an isolated public demo, configurable Machines, a visual Arc workflow, approval handling, durable jobs, and restart recovery paths. The public demo uses Mock Runtime and seeded example history. It cannot connect credentials and has a seven-day visitor lifecycle.
OpenAI and Anthropic adapters exist for registered workspaces. Real-provider acceptance remains pending because no AI provider key has been supplied. Mock execution demonstrates application behavior, not live model reliability or paid workload economics.
Present limitations
Hosted function execution is bounded to 300 seconds; longer work requires a dedicated worker. A deployed target still needs a functioning scheduler or worker to make progress. Runtime and tool throttles are local to a process, while login and demo provisioning limits share PostgreSQL state across instances.
Tools requiring an isolated code sandbox remain unavailable. Provider pricing is configured in a registry and may differ from actual billing. Recovery deliberately stops ambiguous external actions. These are practical boundaries for evaluation and deployment, rather than completed guarantees of production readiness.
Validation roadmap
The immediate validation path begins with a short real-provider task, then a deterministic tool call, and then an approval-controlled action. Each stage should confirm the requested provider, recorded usage, outcome, and behavior under interruption.
The next product questions concern repeated use: which workflows users trust, how much review they need, whether outcomes meet clear criteria, and what execution and support cost. Longer-running work, broader integrations, and additional operational controls should follow evidence from those evaluations.