Skip to content
Hek846Public

About

Standalone agent device firmware for the M5Stack Tab5 (ESP32-P4): on-device chat UI, tool-calling agent loop, persistent memory, and a MicroPython runtime for apps the agent writes itself.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

tab5-claw

tab5-claw is now a standalone Tab5 agent device (M5Stack Tab5, ESP32-P4). The full agent loop runs on-device over its own Wi-Fi: chat UI, tool loop, Moonshot HTTPS calls, on-device memory, and MicroPython app execution.

The Mac host path still exists, but it is now a legacy/tethered debug mode.

flowchart TB
    subgraph tab5["Tab5 device (ESP32-P4, ESP-IDF v5.4)"]
        direction TB
        shell["Shell v2 UI (LVGL)<br/>status bar · launcher/dock · home bar<br/>swipe-back · error banner"]
        chat["Chat app<br/>text + voice input"]
        settings["Settings app"]
        files["Files app<br/>/sdcard/media + /sdcard/agent"]
        agent["Agent loop<br/>in-process tool calling"]
        tools["Tool registry (38 entries)<br/>apps + planning/design · memory · settings · hw_diag<br/>SD sandbox · RS485 · mic/TTS"]
        mp["MicroPython runtime<br/>/flash/apps · crash containment<br/>force-kill (home bar long-press)"]
        mem["Memory<br/>/flash/memory · MEMORY.md<br/>daily notes · retrieval + merge"]
        sd["SD card<br/>/sdcard/media + /sdcard/agent<br/>FAT32 format via Settings only"]
        llm["llm_client (HTTPS)"]
        stt["stt_client (voice)"]
        c6["ESP32-C6 co-processor<br/>Wi-Fi over SDIO (esp_hosted)"]

        shell --- chat & settings & files
        files --> sd
        chat --> agent
        agent --> tools
        tools --> mp
        tools --> sd
        agent --> mem
        agent --> llm
        chat --> stt
        llm & stt --> c6
    end

    cloud["LLM API<br/>(Moonshot, NVS-configurable)"]
    asr["FunASR (Whisper-class ASR)<br/>LazyCat pod on Jetson<br/>LAN :9977"]
    mac["Mac host<br/>(legacy, retired)"]

    c6 -->|HTTPS| cloud
    c6 -->|HTTP, LAN| asr
    tab5 -.->|USB serial JSONL| mac
Loading

Quick start (standalone-first)

Required IDF patch (do this BEFORE idf.py build)

WARNING: a stock ESP-IDF v5.4 builds fine but aborts at runtime.

The firmware REQUIRES a one-hunk local patch to ESP-IDF v5.4: docs/idf-patches/esp-idf-v5.4-xip-psram-tlsp.patch. It adds the EXTRAM range to esp_ptr_executable() in components/esp_hw_support/esp_memory_utils.c, gated on CONFIG_SPIRAM_FETCH_INSTRUCTIONS.

Why: with XIP-from-PSRAM, a stock IDF aborts at runtime in vTaskDelete's TLSP check for any self-deleting task that touched pthread TLS — that means all TLS/HTTP workers and MicroPython app stop/force-kill. The build succeeds; the crash only shows up at runtime, which is why firmware/CMakeLists.txt fails the configure step when the patch is missing.

Apply it once, and re-apply after any IDF update or re-clone:

cd $IDF_PATH && git apply /path/to/repo/docs/idf-patches/esp-idf-v5.4-xip-psram-tlsp.patch
# 1. Flash firmware (ESP-IDF v5.4 + required patch above)
cd firmware
source ~/esp/esp-idf/export.sh
idf.py set-target esp32p4 && idf.py build
idf.py -p /dev/cu.usbmodem1101 flash
# 2. Open a serial JSONL console at 115200 and provision once:
{"type":"wifi_config","ssid":"<YOUR_SSID>","pass":"<YOUR_PASS>"}
{"type":"llm_config","base_url":"https://api.moonshot.cn/v1","model":"kimi-k2.7-code","api_key":"<YOUR_API_KEY>","standalone":1}
{"type":"wifi_status"}
# 3. Use the device directly from the touchscreen (no Mac runtime required).
#    Chat, launch apps, and change settings on-device.

Legacy note: the former Mac-host bridge has been retired. The device runs standalone; the serial JSONL channel below is kept for debugging, provisioning, and regression tests.

Serial protocol (JSONL, one object per line)

This channel is now primarily for debugging, provisioning, regression tests, and legacy tethered mode.

direction message purpose
host→dev {"type":"hello"} request identity; reply includes fw, mac, tools[]
host→dev {"type":"ping"} liveness check (pong)
host→dev {"type":"status","text":..} set title/status text (ack)
dev→host {"type":"heartbeat",..} periodic uptime/status heartbeat
dev→host {"type":"chat_request","id":N,"text":..} legacy bridged chat request
host→dev {"type":"chat_status","id":N,"state":..} update temporary chat state
host→dev {"type":"chat_response","id":N,"text":..} bridged reply bubble, then chat_done
host→dev `{"type":"chat_session","op":"list new
host→dev {"type":"tool_call","id":N,"name":..,"args":{}} invoke registered device tool (tool_result)
host→dev {"type":"app_run","name":"<app>"} / {"type":"app_stop"} / {"type":"app_unskip","name":"<app>"} runtime app control (app_result / app_unskip_result). app_stop sets ok true only when a running app actually stopped; otherwise err is not_running or still_running
host→dev {"type":"fs_write","path":..,"text":..} atomic whole-file UTF-8 write under /flash (temp + rename; fs_result); migration tooling may add offset/final and encoding:"base64" for restart-from-zero binary-safe writes
host→dev {"type":"fs_read","path":..,"offset"?:N} read under /flash (fs_result); offset text pages are ≤4000 bytes and never split UTF-8; migration may request encoding:"base64" for ≤2400-byte binary pages; legacy no-offset reply is unchanged
host→dev {"type":"fs_list","path":..} list a /flash directory (fs_result)
host→dev {"type":"fs_mkdir","path":..} create one validated directory under /flash; developer migration path only (fs_result)
host→dev {"type":"llm_config",...} store/update LLM config in NVS (llm_config_ack)
host→dev {"type":"llm_config_get"} return model, base URL, standalone mode, and whether a key is configured (llm_config_status; key never returned)
host→dev {"type":"llm_test"} single-shot LLM latency probe (llm_test_result); see below
host→dev {"type":"wifi_config","ssid":..,"pass":..} store Wi-Fi creds in NVS (wifi_ack)
host→dev {"type":"wifi_status"} query current Wi-Fi state ({"type":"wifi",...})
host→dev `{"type":"mode","standalone":0 1}`
host→dev {"type":"ask","text":..} / {"type":"chat_inject","text":..} test injection hooks; v1 success is ack with request_id of the accepted request. Busy / rejected inject replies error chat_busy (never an ack carrying a previous round's id)
host→dev `{"type":"chat_test_mode","enable":true false}`
host→dev {"type":"chat_cancel"} cancel in-flight LLM request (chat_cancel_ack)
host→dev {"type":"chat_state"} query agent state/session/elapsed; additive done_kind (done|error|waiting_user|cancelled), last_done_request_id, last_reply after a finished request
host→dev {"type":"app_plan_get"} read the current session's pending app plan and ordered questions (app_plan_state)
host→dev {"type":"app_plan_approve","confirm":true,"revision":N,"session_id":N,"answers":[...]} explicitly approve the exact pending plan over the local serial developer link and submit its build continuation (app_plan_approve_ack)
host→dev {"type":"chat_session","op":"history","id":N,"max":M} load session history (chat_session_history)
host→dev {"type":"screen_capture","scale":N} framebuffer readout (screen_capture_begin + N×screen_capture_chunk + screen_capture_end)
host→dev {"type":"hw_camera_diag","id":N,"capture":bool} developer-only camera probe: sensor detect + optional RGB565 still (hw_camera_diag_result); deliberately NOT an agent tool
host→dev `{"type":"touch_inject","action":"down up
host→dev `{"type":"host_takeover","active":true false,"label"?:".."}`

Serial plan approval preserves the same firmware trust boundary as the touch flow: the host must first read the pending plan, return its exact revision and session, provide one non-empty answer per question, and set confirm:true. Stale plans, mismatched answer counts, and approvals while Chat is busy are rejected.

sequenceDiagram
    participant Host as Developer host
    participant Serial as Serial JSONL
    participant Plan as App plan policy
    participant Agent as On-device Agent

    Host->>Serial: app_plan_get
    Serial->>Plan: copy pending plan for current session
    Plan-->>Host: app_plan_state (revision, questions, options)
    Host->>Serial: app_plan_approve (confirm, identity, answers)
    Serial->>Plan: validate exact pending plan and grant create permission
    alt stale, malformed, or busy
        Serial-->>Host: correlated error
    else approved
        Serial->>Agent: submit generated approval prompt
        Serial-->>Host: app_plan_approve_ack (request_id)
        Agent->>Agent: continue app_write and verification
    end
Loading

Host-takeover indicator

When the developer host drives the Tab5 over the serial dev-plane, a compact amber HOST pill appears at the left of the status-bar cluster (beside the Wi-Fi/clock/app-status items) so a bystander can tell at a glance that the computer is in control. It borrows width from the title using the same shrink rule as the Agent progress label, so it never clips the title or pushes the Wi-Fi/clock cluster off-screen.

Detection is activity-based, not a single trust-me flag the host might crash before clearing (firmware/main/host_takeover_policy.h, host-tested by tests/test_host_takeover_policy.c):

  • Any control/mutation message (chat_inject, ask, app_plan_approve, app_run/app_stop/app_unskip, tool_call, touch_inject, fs_write, fs_mkdir, safe_mode, open_settings, status, chat_status, chat_cancel, chat_test_mode, chat_session, memory_update, mem_consolidate, wifi_config, llm_config) refreshes a monotonic "last host activity" timestamp and latches the badge on.
  • Read-only / observation messages are deliberately passive and never light the badge: hello, ping, chat_state, wifi_status, llm_config_get, llm_test, app_plan_get, fs_read, fs_list, hw_rtc_diag, hw_camera_diag, and screen_capture. This is what lets the badge clear during long polling — chat_drive.wait_idle polls chat_state with backoff capped at 8 s — and lets dev tooling screenshot the screen to verify the badge cleared without the capture re-arming it.
  • A permanent LVGL sweep timer clears the badge once the host has been quiet for HOST_TAKEOVER_IDLE_MS (12 s), longer than the 8 s poll cadence so a pure poll stream can never keep it alive.
  • The optional explicit host_takeover message lets a host announce intent: active:true is a labeled activity ping (still subject to the 12 s idle safety net so a crashed host cannot pin it forever), active:false forces the badge off immediately.
stateDiagram-v2
    [*] --> Idle
    Idle --> Driving : control/mutation msg\nor host_takeover active:true
    Driving --> Driving : another control msg\n(refresh timestamp)
    Driving --> Idle : 12 s since last activity\n(LVGL sweep timer)
    Driving --> Idle : host_takeover active:false
    note right of Idle
        Passive msgs (chat_state poll,
        screen_capture, reads) do NOT
        change state
    end note
Loading

LLM probes & telemetry (llm_test + Settings → AI Services → AI Tests)

Developer-only checks shared by the serial llm_test message and the Settings → AI Services → AI Tests sub-page. The Response speed row used to live on the main AI Services list; it now lives on the AI Tests page alongside two more probes and last-task/token telemetry. One busy gate (llm_test_try_begin/end/busy) covers every trigger, so a second request while any test is in flight returns immediately with err: "busy" (serial) or an "AI test busy" / "Endpoint busy" status (UI) — never a fake millisecond value. Each probe is a single HTTPS chat attempt with a fixed short fixture prompt — no 429 retries — and the wall-clock stamp starts only after the LLM client lock is acquired (3 s bounded wait; lock miss is also reported as "busy" with ms: 0).

The AI Tests page (app_settings_ai_tests.c) shows:

  • Last Agent Task banner — retained terminal snapshot captured at agent finish (outcome, tool steps, tool ok/fail, elapsed); survives the next request's worklog wipe.
  • Probes — Response speed (tools-off, llm_test_measure), Tools-on latency (a fixed 2-tool no-side-effect subset — get_status + app_list — wrapped in the same OpenAI {type:"function"} shape as the agent manifest via tool_registry_wrap_openai), Tool correctness (asks for a get_status tool call and checks the reply carries a legal tool_calls entry — the tool is never executed; Pass/Fail + ms), and Run all (sequential tools-off → tools-on → correctness; continues past a busy/HTTP failure, stops only on OOM).
  • Token usage — last prompt/completion/total tokens from llm_client_last_usage(), persisted in NVS (claw_aitst) so it survives a reboot.

The serial llm_test message is unchanged (still returns llm_test_result).

sequenceDiagram
    participant Host as Host / Settings UI
    participant Gate as llm_test busy gate
    participant Worker as TLS worker (≥12KB)
    participant LLM as llm_chat_once
    participant Out as llm_test_result / AI Tests row

    Host->>Gate: try_begin
    alt gate already held
        Gate-->>Host: busy (serial err / UI status)
    else acquired
        Host->>Worker: start measure
        Worker->>LLM: fixture prompt (lock ≤3s)
        alt lock timeout / busy
            LLM-->>Worker: CLAW_ERR_LLM_BUSY
            Worker-->>Out: ok:false err:"busy" ms:0
        else request completes
            LLM-->>Worker: reply + elapsed ms
            Worker-->>Out: ok / ms / text (or err)
        end
        Worker->>Gate: end
    end
Loading

Reply shape (llm_test_result):

field meaning
ok true on a completed probe; false on busy or failure
ms request wall ms after lock acquire; 0 when no request left the device
text model reply content (success only)
err "busy" when the shared gate or LLM lock rejects; otherwise esp_err / …/http_N

Memory

  • Memory is now device-owned under /flash/memory:
    • MEMORY.md (long-lived profile/facts)
    • daily notes + retrieval over those notes
    • LLM-driven consolidation back into MEMORY.md
  • The former host memory bridge is retired.
  • Device memory tools are exposed through the on-device agent tool loop.

Agent task UX

  • Agent work is not a modal screen lock. A Home Bar tap or upward gesture navigates to the launcher while the request continues, so the user can open another app. A non-interactive pulsing yellow edge remains visible over the shell as the global work indicator; it reports activity but does not consume touch input or pin the current screen.
  • Stopping is deliberately destructive. Long-pressing the Home Bar while Agent work is active opens a Keep working / Stop task confirmation. Only the confirmed stop cancels the request (and stops a running generated app, if one is active). The live Chat progress cell's Stop button uses the same confirmation. Ordinary navigation never cancels Agent work.
  • Progress belongs to the originating Chat session. Its live cell shows phase, tool-step count and the latest factual milestone; tapping it expands a bounded timestamped worklog of plan/tool outcomes. Tool arguments, source and payload bodies are not copied into that log. A completed request keeps a compact expandable history summary before its final reply.
  • Background outcomes do not steal focus. Completion, failure and input-needed outcomes show a tappable banner when the originating session is not already visible. Tapping routes to that session; the shell never opens Chat or switches sessions on its own.

Hardware API (agent-facing)

  • SD sandbox — sole registration in sd_sandbox.c: the agent sees only /sdcard/agent via sd_agent_list / sd_agent_read / sd_agent_write. Name policy is sd_sandbox_path (shared with hw_sd_agent_*). Writes are create-only (never overwrite or delete), ≤4KB, and require confirm:true. Long filenames enabled (CONFIG_FATFS_LFN_HEAP); txt/md/json/csv extensions.
  • Files / Media (local UI) — the launcher Files tile (pin name agent_files) is a human-only browser of /sdcard/media and /sdcard/agent. It is not an agent tool and does not change create-only sd_agent_write. Path policy is files_path_policy (host-tested): no traversal, no hidden/system names, no cross-root moves, and text/JPEG previews are size-capped (4KB / 2MB). Rename stays inside one root; delete requires an explicit confirm. The /sdcard/media and /sdcard/agent roots themselves are never renameable or deletable from Files — Camera, Recorder, and the agent sandbox depend on those directories existing after format, and removing a root would break those contracts and risk orphaned or lost data; children inside each root remain locally manageable. Missing or remounted cards are shown honestly. Identity notes (Soul / About you / Memory) stay on /flash and are opened from this app so the old Agent files tile is not duplicated.
  • SD format — Settings → SD browser only, behind a two-step "Erase all SD data?" confirmation; formats FAT32. No agent or serial format path; the BSP keeps format_if_mount_failed = false.
  • Diagnostics — hw_diag_rtc/power/switches/i2c_scan/imu/storage/ audio_volume, plus whitelisted hw_diag_gpio.
  • Speech — mic_listen (1-15 s → LAN STT) and tts_speak (1-300 chars → Kokoro TTS); both require consent:true on every call and are busy-gated against Chat voice.
  • Configure from computer — Settings can open a five-minute, LAN-only HTTP setup page protected by a one-use six-digit code. The page uses plaintext HTTP, so use it only on a trusted LAN. Secrets are sent only in capped POST bodies, remain pending until explicit confirmation on the device, and are discarded when the session closes or expires.
  • RS485 — hw_rs485_txrx only: enable:true per call, 1-64 bytes hex, 20-1000 ms timeout, ~200 ms rate limit. No GPIO/power/bus-scan access.
  • Camera — camera_broker.c is the sole owner of SC202CS access, exclusive leases, cancellation and the title-bar privacy indicator. The system Camera app keeps continuous preview and mirror behind a preview lease. It samples BMI270 orientation off the LVGL task, locks it at shutter time, and saves portrait or landscape JPEG pixels to /sdcard/media; the viewer and thumbnail gallery preserve aspect ratio. Agent camera_capture takes exactly one portrait photo and requires consent:true on every call. Generated apps use claw.camera.request_permission() followed by foreground-only claw.camera.capture(); the grant lasts only for that app run and the API returns metadata, never raw pixels or /dev/video0. The serial hw_camera_diag probe remains developer-only but also goes through the broker.
  • Level — standalone system app with a 50Hz worker-backed two-axis bubble, X/Y angles and live IMU health classification. Quick calibration samples the stationary BMI270 for three seconds, rejects missing, moving, implausible or non-level input, and stores the accepted zero in NVS. BMI270 provides pitch/roll only; the app does not claim compass yaw.
  • Clock — built-in local time, stopwatch, countdown timer and up to eight persistent alarms. Alarm edits are committed from a worker, and due alarms use a visible system alert plus best-effort shared-audio beep. Alarms run only while firmware/LVGL is active; there is no RX8130 powered-off wake path.
  • Recorder — standalone system app for voice memos. Recording streams 16 kHz mono 16-bit PCM from the mic straight to a .wav.tmp file under /sdcard/media/recordings (never staged in RAM, capped at two hours), then patches the WAV header, fsyncs and renames; an interrupted capture is recovered into a playable memo on the next list refresh. Playback is chunked streaming with mono-to-stereo upmix. Memos support rename (policy-checked base names, never overwrite) and two-step delete. Recording and playback go through the shared hw_mic/hw_audio busy gates, so the Recorder and Chat voice / mic_listen / tts_speak / Settings audio tests mutually refuse while any of them holds the audio path. Its streaming and memo workers also reserve the shared SD lock for their complete operation because FATFS file locking is disabled. Recorder is local UI only — not exposed to agent tools or the serial protocol.
  • UI introspection and geometry assertions — one read-only tool, ui_inspect, dispatches on action. tree pages widgets with stable id/name/type, rendered geometry, text and retained canvas draw ops. Tree JSON is labelled coord_space:"screen"; canvas_origin is the claw.ui local (0,0) measured in that same screen space (typically near 18,82: status bar plus content pad). Hardware touch and touch_inject use screen pixels; generated claw apps keep canvas-local coordinates — the app coordinate system is not changed. element returns one widget by id or name; check reports label/button text clipping, content-bound violations, overlaps covering at least one quarter of the smaller widget, and button/slider targets below THEME_TOUCH_MIN. Results are capped at 4KB and tree pages at 24 widgets. assert_center compares one target with the app content rectangle or another named/id widget using absolute rendered boxes. Axis defaults to both, tolerance to 4px (maximum 64), and target measurement may be object_box or text_object_box; the result includes both rectangles and centers, dx/dy, per-axis verdicts and pass. text_object_box is available only where LVGL exposes a text object (label, button label or textarea label), and a reference widget is always measured by its object box. These actions do not prove optical glyph balance, pixels drawn inside a canvas, color contrast, or a human interaction; those still need the corresponding review. Apps adjust and remeasure through claw.ui.set_pos / set_size, stable name= arguments, and claw.ui.text_width(text, role) when a title-role label is being measured; the one-argument text_width form intentionally keeps body-font metrics.
flowchart LR
  hwTouch[Hardware touch] --> screenPx[Absolute screen pixels]
  touchInject[touch_inject] --> screenPx
  inspectTree["ui_inspect tree x,y"] --> screenPx
  clawUi["claw.ui widgets"] --> canvasLocal[Canvas-local]
  canvasLocal -->|"plus canvas_origin"| screenPx
Loading
  • Targeted source navigation — app_read defaults to outline, a bounded best-effort structure summary of imports, globals, simple claw.ui.<constructor> assignments, functions/classes and the claw.run entry point. Explicit actions are read_range (1-based inclusive, at most 200 lines/3072 bytes), read_symbol (one exact top-level function/class, with distinct absent/ambiguous errors), literal case-sensitive search (up to 12 matching lines with 0–2 context lines), and full. Every successful source read reports the exact-byte revision as fnv1a32:<8-hex>, byte/line counts, and malformed/truncation/best-effort flags. The outline scanner is intentionally not a full Python parser.
  • Revision-guarded app editing — replacing an existing app with app_write and every app_edit require the latest expected_revision from app_read. The current file is rehashed under the shared mutation lock and a missing, malformed or stale revision is rejected before any write; a stale reply includes the current revision and asks for a fresh read. A new app_write omits the revision and still passes the plan gate. Successful mutations are atomic and return their new revision. app_edit then replaces one unique raw-byte anchor; zero or multiple matches are distinct errors. Line numbers are navigation only—there is deliberately no line-number mutation API. The anchor logic remains shared with the notes edit path.
  • Generated-app design system — the grouped app_design tool exposes catalog, paged tokens (maximum 12 per call), template, guide, and checks actions. Its current catalog contains 18 semantic tokens, 10 component/screen templates (page_title, card, three button roles, list_row, form, empty_state, confirmation, game_start_menu), five compact guide topics and eight deterministic check policies. theme.h remains the concrete source of truth for background/surface/primary, text/muted/status/border colors, body/title type roles, spacing, radii, minimum touch size and transition timing. Game art and accent palettes may vary, while controls and semantic structure stay consistent. The catalog is guidance, not rendered-layout evidence; fetch the nearest template and only the relevant token/guide pages, then verify the running UI.
  • In-app theme lookup — generated Python calls claw.ui.theme("color.primary") (or another catalog token) to resolve the same value used by the shell. Color, spacing, radius, touch and timing tokens return integers; type.body / type.title return their semantic role strings. claw.ui.set_text_role(id, claw.ui.theme("type.title")) applies one of those roles to the actual text object of a label, button or textarea; another role, an unknown widget id or a non-text widget raises ValueError rather than falling back silently. claw.ui.theme itself does not apply a style, and claw.ui.set_mono remains the separate explicit monospace API.
  • Deterministic app verification gate — each Agent request derives proof only from parsed arguments and top-level results of the real app tools, not from model prose. For every successful app_write/app_edit, it requires a stop when old code was still running, a hidden non-foreground start of the latest mutation that reaches READY for the exact app/generation/source, ui_inspect({action:"check"}) for that run with zero findings and at least one checked widget, and app_last_crash with an empty record for that app/run. The same runtime identity must still be READY when the final answer is gated. If the request contains the case-insensitive English whole word center, centre, centered or centred, or the exact Chinese text 居中, it additionally requires a failing, fully specified named assert_center measurement before mutation followed by the identical passing measurement after restart. This causal pair binds target, reference, axis, tolerance and text-object measurement, rather than accepting an unrelated already-centered widget. A later mutation invalidates earlier start/UI/center/ crash evidence; failed/stale mutations, wrong-app results, malformed result shapes, multi-app work and the bounded eight-app tracking capacity fail closed. If a prose reply is missing one safe runtime check (app_stop, exact app_start, clean ui_inspect, or app_last_crash) and tool budget remains, the gate feeds back the exact app and one missing action and continues the same request. It stops honestly on an unchanged missing action, near the step limit, or for mutations, UI fixes, causal-center baselines and human actions. Otherwise the final model answer is replaced with Not verified plus the exact next check instead of preserving an unsupported success claim. This gate is scoped to the current request and its basic runtime/UI checks; design-specific contrast, optical glyph balance and human gestures remain separate evidence.
  • App planning gate — creating a new /flash/apps/<name>.py requires an interactive plan confirmed in Chat. The grouped app_plan_request tool supplies 4–6 questions with 2–4 choices; the card handles selection, optional free text, Back/Next, review, cancellation and confirmation locally without an LLM call per tap. app_write enforces the matching session and app name in firmware and returns plan_required otherwise. Existing-app repairs remain ungated. An approved plan is saved under /flash/data/app_plans/ after the first successful write. Repeating app_plan_request for the already approved app/session returns the stable plan_already_approved error and tells the model to call app_write; it does not replace the approval or present another card.
  • Notes — one tool, notes, with list / read / append / edit / delete over a fixed /flash/notes/ directory on internal flash (deliberately not the SD card, which can be absent). append is the point: the SD sandbox's create-only policy is what forced the on-device agent to fragment one note into many files. 16KB per note, 4KB per read. The developer side pulls with python3 tools/dev_cycle.py notes pull into a local docs/device-notes/ directory. delete is the one place the agent may destroy its own data, and the divergence from the SD sandbox is deliberate. The sandbox is create-only because it holds the user's card; /flash/notes is the agent's own outbox to the developer, and a channel you cannot correct is a channel you work around — which is exactly how the fragmentation happened. It is bounded to match: confirm:true per call (stateless, like mic/TTS consent), one named file per call, the same name policy as append, no wildcards, no recursion, and no directory removal. A missing note is not_found, never a silent success.
  • Periodic apps — app_schedule combines set/clear/list for at most 16 persistent background schedules (60 seconds to one year). Every occurrence uses the normal fresh MicroPython VM lifecycle. Per-app claw.state persists scalar values with a 4KB cap. Scheduled runs cannot draw, wait in claw.run(), or call the network-backed claw.agent.ask(); a due run is skipped while the device is busy. Automations is the built-in system app (app_automations.c, launcher tile + firmware/main/automations_policy.[ch]) that manages this same periodic_apps.c registry — one scheduler, one persistence/locking path, no second store. It lists every schedule with app, interval, enabled state, next/last run, and the last honest result or skip reason (skipped_busy/disabled_crashes/disabled_file_error/…), and offers enable/disable, edit-interval (a preset roller still bound by the 60s..1y policy), and delete with an explicit two-step confirmation. Run-now only starts when periodic_apps_runtime_busy() — the exact gate scheduler_task itself checks (MicroPython running, an agent request in flight, or the UI marked busy) — is clear; otherwise the button shows the honest reason ("busy" or "disabled") instead of silently no-op'ing. Settings keeps a single "Automations" row that opens this app rather than a second, independently-mutating list.
  • Tool registry cap is 40 (tool_registry.c), 38 currently registered (pinned by tests/test_prompt_consistency.py). The real cost of a tool is not RAM: every tool's name, description and schema is serialized into every LLM request. Group related capability behind one tool with an action field rather than adding tools freely.

Known limitations (status as of 2026-08-06):

  • RTC — fixed and hardware-verified at d9f9718. The legacy writer used numeric weekday and zero-based month and never cleared VLF; separately, Tab5 backup switching and charging were disabled (Control1 0x00 on cold boot). The boot guard rejecting RTC years outside 2024-2099 stays as defence.
  • STT — verified on device (one 2 s mic_listen → LAN STT in 4182 ms).
  • TTS — not verified. tts_speak fails with ESP_ERR_HTTP_CONNECT before synthesis. The gate was explicitly deferred on 2026-08-04; that is an acceptance exception, not a pass.
  • Real RS485-device validation, camera cancellation during the narrow capture window, and vision-model flow — not verified.
  • Repeated SD format resets the device (task WDT → SW_CPU_RESET); root cause unproven. A boot-scoped latch disables Format after one success until reboot — a mitigation, not a fix, and itself unexercised on hardware.

Hardware bring-up notes (hard-won; do not lose)

  • Tab5 panel/touch is ST7123 (in-cell). The BSP auto-detect can fail before the panel is powered; our vendored BSP defaults to ST7123 (the ST7703/ILI9881 path hangs the DSI FIFO on this hardware).
  • The vendored esp_lcd_st7123.c once contained a debug hack that skipped the tail of the init command table — including SLPOUT (0x11) and DISPON (0x29) — producing a perfectly "working" black screen. If the display ever goes dark again with clean logs, check the init loop sends all 27 commands. (The read-id probe skip is intentional and must stay: DCS read 0x04 stalls DSI.)
  • Touch needs the PI4IOE I/O expander init + bsp_reset_tp() before bsp_display_start() (power sequence from the factory demo), or it NACKs.
  • LVGL tear-avoidance must stay off: the BSP's DPI path allocates a single hardware framebuffer.
  • CONFIG_ESP_MAIN_TASK_STACK_SIZE=8192; display boot runs in a task pinned to CPU1 so serial stays responsive.
  • exFAT drive: run find . -name '._*' -delete if the build chokes on AppleDouble files inside managed_components.
  • Moonshot free tier is rate-limited (~3 req/min); the agent waits out 429s.

Wi-Fi (ESP32-C6 co-processor)

Working since fw 0.4.0: esp_hosted 1.4.0 + esp_wifi_remote 0.8.5 over SDIO. Credentials: defaults in main/wifi_creds.h, override via serial {"type":"wifi_config","ssid":"...","pass":"..."} (stored in NVS); query with {"type":"wifi_status"}. On boot the device connects, runs a plain-HTTP reachability probe, and reports {"type":"wifi","state":"online",...}.

Bring-up traps (all fixed in sdkconfig.defaults, keep them):

  • ESP_HOSTED_SDIO_PIN_* and ESP_HOSTED_IDF_SLAVE_TARGET are computed Kconfig symbols — setting them in defaults silently does nothing. The settable ones are CONFIG_ESP_HOSTED_SDIO_PRIV_PIN_D1_4BIT_BUS=10 and CONFIG_SLAVE_IDF_TARGET_ESP32C6=y.
  • Tab5 SDIO pin map: CLK 12, CMD 13, D0 11, D1 10, D2 9, D3 8, C6 reset 15. With the wrong D1 the SDIO probe fails (failed to get CIS data) and esp-hosted assert-loops the board.
  • The Wi-Fi/TLS stack needs a custom 16 MiB layout: two 6 MiB A/B app slots plus 3.875 MiB LittleFS storage. USB flashing remains the recovery path.

Test & acceptance tools

Host unit tests need no hardware:

make test   # native C test executables, then every Python test_*.py

Everything below drives a real device over the serial link. They share tools/devicelink.py; per-script differences (read chunk size, message/log capture, per-line hooks) are constructor arguments, so do not fork it again.

Daily driver: tools/dev_cycle.py

One thin entry point over the devicelink primitives for the personal one-device loop. Omitting --port always means exact hello-identity discovery — no port is ever hard-coded or assumed. Flashing always uses the resolved explicit port and never erases. The default cycle smoke path sends no chat, touch, Wi-Fi/API/config, session, or network-LLM traffic; the interaction helpers (status, tap, chat-test) only run when invoked explicitly. chat-test uses the deterministic echo test mode in one throwaway session, restores the original session, and deletes only the session id it created — it never calls a network LLM. get/put/list stay inside the /flash sandbox; put accepts at most 32 KiB of valid UTF-8 text, and any transfer failure restarts the whole operation.

python3 tools/dev_cycle.py info                 # discover + identity/caps
python3 tools/dev_cycle.py cycle                # build → no-erase flash →
                                                # reconnect → ping/heartbeat-status
                                                # → bounded log → run.json
python3 tools/dev_cycle.py cycle --skip-build --skip-flash   # attach only
python3 tools/dev_cycle.py cycle --probe        # + /flash 5 KiB UTF-8
                                                # round-trip fixture check
python3 tools/dev_cycle.py cycle --capture      # validate screenshot stream
                                                # in memory; save metadata only
python3 tools/dev_cycle.py cycle --screenshot   # export pixels only after the
                                                # screen is confirmed non-sensitive
python3 tools/dev_cycle.py get /flash/data/devcycle_probe.txt --out probe.txt
python3 tools/dev_cycle.py put /flash/data/fixture.txt ./fixture.txt
python3 tools/dev_cycle.py list /flash/data
python3 tools/dev_cycle.py capture --out shot.png   # explicit screenshot
python3 tools/dev_cycle.py status "building"    # explicit: changes UI text
python3 tools/dev_cycle.py tap 360 640          # explicit: injects one tap
python3 tools/dev_cycle.py chat-test            # explicit: echo round-trip
python3 tools/dev_cycle.py plan-show             # inspect pending plan/questions
python3 tools/dev_cycle.py plan-approve --confirm \
  --answer "4x4 board" --answer "Swipe controls" \
  --answer "Classic colors" --answer "Keep best score" --wait-idle

cycle evidence lands in agent-work/reports/devcycle-<timestamp>/ (content-free bounded serial summary, step list, final pass/fail run.json); raw serial lines are never persisted, so credentials, config, chat content, and environment values are not collected.

# Fault injection: network drop mid-chat, force-kill inside claw.agent.ask(),
# and 100 rounds of app_stop racing a force-kill.
python3 tools/gate_fault_inject.py /dev/cu.usbmodem1101 --scenario all
python3 tools/gate_fault_inject.py /dev/cu.usbmodem1101 --scenario race --rounds 100

# Regression test for cross-session contamination during a streamed reply
python3 tools/gate_e19_session_switch.py /dev/cu.usbmodem1101 20

# Render-bound benchmark (session-switch latency; baseline ~22ms)
python3 tools/render_bench.py /dev/cu.usbmodem1101 20

# Hunt the taskLVGL wedge signature (heartbeat alive while task_wdt fires)
python3 tools/wedge_hunt.py /dev/cu.usbmodem1101 6

# Gate A: chat state-machine acceptance (deterministic echo mode)
python3 tools/gate_a_acceptance.py /dev/cu.usbmodem1101 --rounds 3

# Gate A: real LLM path
python3 tools/gate_a_acceptance.py /dev/cu.usbmodem1101 --rounds 3 --qwen

# Gate B: screen capture + touch injection verification
python3 tools/gate_b_capture.py /dev/cu.usbmodem1101 verify

# Gate B: manual capture / tap / swipe
python3 tools/gate_b_capture.py /dev/cu.usbmodem1101 capture --out shot.png
python3 tools/gate_b_capture.py /dev/cu.usbmodem1101 tap 360 640
python3 tools/gate_b_capture.py /dev/cu.usbmodem1101 swipe 360 900 360 300

# Gate C: full autonomous build → flash → verify loop
python3 tools/gate_c_loop.py /dev/cu.usbmodem1101 --rounds 3
python3 tools/gate_c_loop.py /dev/cu.usbmodem1101 --rounds 3 --qwen
python3 tools/gate_c_loop.py /dev/cu.usbmodem1101 --skip-build  # skip build if already flashed

Roadmap

  1. Phase 2 (DONE): untethered LLM + on-device memory/tools (llm_client + NVS config, layered /flash/memory, in-process tool loop, self-evolving MicroPython app workflow).
  2. Shell v2 (DONE): status/home/launcher+dock, chat sessions, and Settings/About shipped in current firmware rounds.
  3. Phase 1 (DONE, 2026-08-01): reliable chat workbench (session-bound state machine, real cancellation, crash recovery, deterministic test mode), autonomous visual QA (screen capture + touch injection), autonomous repair loop (build → flash → verify). See DEVLOG.md 2026-08-01 for the gates.
  4. Voice (PARTIAL): consent-gated agent speech tools landed 2026-08-03 (mic_listen → LAN STT, tts_speak → Kokoro TTS). Network end-to-end validation still pending.
  5. Camera (PARTIAL): SC202CS continuous preview, PPA rotation/scaling, mirror, BMI270-oriented portrait/landscape SD capture and thumbnail gallery ship in the system Camera app. Shared broker, consent-gated Agent capture and foreground-only generated-app API are implemented and device-verified; in-flight cancellation and the vision-model flow remain.
  6. More hardware tools (PARTIAL): hw_diag_* suite (INA226 power telemetry, IMU, I2C scan, GPIO whitelist), SD sandbox, and restricted hw_rs485_txrx landed 2026-08-03. Still open: RTC alarms (RTC time itself is currently wrong — known issue), IMU gestures, real RS485 device validation.
  7. Self-update (DONE): signed OTA A/B installer with on-device check/confirm/progress, stable public manifest feed and health-gated rollback. Network install, slot switching, data preservation and forced rollback passed.

About

Standalone agent device firmware for the M5Stack Tab5 (ESP32-P4): on-device chat UI, tool-calling agent loop, persistent memory, and a MicroPython runtime for apps the agent writes itself.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages