Infrastructure Upgrade with the Agents API
My NAS and server aged quite well—both were commissioned between 2017 and 2019—but they needed a refresh, so I chose close successor models to keep up with growing performance demands.
nas.lan moved from a QNAP TS-431XeU to a TS-432XeU, and the board in server.lan moved from an ASRock J5005-ITX to an ASRock N100M. At the same time, I turned the OpenClaw inspection routine that originated on bot.lan into a Server-owned service that can diagnose and repair the three core hosts through OpenAI’s new Agents API.
Both projects support the same operational goal: the agent runs on reliable hardware and retains a recovery path when the normal network fails.
Two hardware swaps
The NAS replacement was the simpler migration. QNAP lists the TS-431XeU to TS-432XeU move as a supported QTS system migration. I updated the empty target first, shut down both units, and transferred all four labeled drives into their original bay numbers. QTS retained the configuration and storage layout; the 8 TB HDD RAID 1 and 400 GB SSD RAID 1 returned healthy.
The new unit replaces the old 32-bit Alpine AL314 and DDR3 platform with a 64-bit Alpine AL324 and DDR4. My unit has 8 GB of RAM. It retains the 10 GbE SFP+ port and adds two 2.5 GbE ports in place of the former two 1 GbE ports (TS-431XeU specifications, TS-432XeU specifications). I updated the Kea reservations for both new network identities, then verified 10 GbE access, 2.5 GbE Wake-on-LAN, and a controlled QTS shutdown.
The server.lan swap required more preparation. The previous J5005-ITX configuration used 16 GB of RAM and a SATA system SSD. The N100M now runs with 32 GB of DDR4-3200 RAM and boots Ubuntu from a 128 GB Patriot P320 NVMe SSD. I kept the 250 GB shared-data SSD, the 500 GB backup SSD, and the PCIe 2.5 GbE adapter. After fixing the switch port, the retained adapter negotiated 2.5 Gbit/s full duplex.
The N100M specification gives it one full-size DDR4 slot for up to 32 GB, a PCIe Gen3 x2 M.2 socket, and separate expansion slots. This let me move boot I/O to NVMe while keeping the 2.5 GbE adapter. I kept the old system SSD disconnected as a rollback option until the new host had passed its controlled reboot and backup reconstruction.
The board also made room for two independent Internet fallbacks: an Intel AX211 Wi-Fi adapter and the existing SIM7600 LTE modem with its separate FTDI control connection. The final acceptance covered a controlled reboot, SMB and backup recovery, both SATA drives, the NVMe boot path, and all 17 containers.
Inspection becomes a service
This continues the operating model from Infrastructure with AI Agents, but removes the inspection job from OpenClaw. A Python coordinator on server.lan now calls the actual /v1/agents/sessions endpoint. The Agents API supplies the managed Codex harness, session lifecycle, context handling, and recovery; my application supplies the tools and execution policy.
This separation is the point. Infrastructure faults can remain unnoticed, and if OpenClaw on bot.lan breaks, it cannot reliably repair its own execution path. The Server coordinator and Agents API provide an external inspection and repair path that does not assume OpenClaw is healthy.
The cloud agent has no shell on my LAN and no daemon runs on gateway.lan or bot.lan. Instead, a small fixed-function gateway executes bounded checks locally on Server or through pinned SSH connections to Gateway and Bot. Inspection always reasons about server.lan first, then gateway.lan, then bot.lan, before it considers the remaining devices.
The normal API path leaves Server over 2.5 GbE through gateway.lan. After sustained failures, the controller enables Wi-Fi and connects directly to Fritz, bypassing Gateway. If that path also fails, it shuts Wi-Fi down and selects LTE. Hysteresis and an independent rollback timer prevent route flapping or a stranded fallback. This gives the coordinator a final outbound path to the Agents API even when the main router path is unavailable.
The workflow is deliberately layered. Before Luna starts, TypeSafe.ai’s Jev handles the initial decision step by turning the collected state into typed, probabilistic decisions. Unlike a general-purpose generative LLM, Jev does not generate prose token by token; its possible outputs are defined in advance, so the coordinator can branch without parsing free-form text. That constrains the output shape, not correctness, so Jev can approve only the strict healthy fast path. Anything incomplete, uncertain, or unhealthy goes to Luna for inspection.
A schema-validated issue then routes routine repairs to Terra and severe or mixed repairs to Sol. Every repair has a stated purpose, a timeout, a separate verification command, and a durable receipt. Runs are single-flight; an uncertain call is never replayed automatically, and an unknown result never becomes a claimed success.
The job runs five times each day in (3h intervalls) and can also be started manually at inspection.server.lan: Here a dashboard exposes the current traffic light, severity counts, connectivity path, model activity, token use, and sanitized history.
The hardware refresh increased capacity, but the more important change is operational: the inspection service has bounded authority, verified repair paths, and two independent ways to stay connected when the normal route disappears.
This now allows the main OpenClaw Agent, that was prior tasked for repairs, to break itself and be fixed up by an secondary independent Agent from the outside.