Persist ride state across app restart and crash (#12) #98

Merged
robert merged 1 commit from area/persist-ride-state into main 2026-09-04 19:23:45 +02:00
Owner

Closes #12.

See the commit message for the full breakdown (what's persisted vs. reset and why, the 30s checkpoint-throttling rationale, how the resume/discard prompt reuses carousel.c's existing ConfirmKind pattern, and exactly what was verified in the emery emulator via a real SIGKILL force-kill/relaunch cycle for both the resume and discard paths).

Not verified: the physical-watch half of the acceptance criterion ("and on the watch") — needs Robert's real emery/gabbro hardware, not done in this sandbox.

Builds on #11 (PR #84, state.c's ride state machine) and #58 (PR #81, page.h's persist-key registry / persist_read_data/persist_write_data pattern).

Closes #12. See the commit message for the full breakdown (what's persisted vs. reset and why, the 30s checkpoint-throttling rationale, how the resume/discard prompt reuses carousel.c's existing ConfirmKind pattern, and exactly what was verified in the emery emulator via a real SIGKILL force-kill/relaunch cycle for both the resume and discard paths). Not verified: the physical-watch half of the acceptance criterion ("and on the watch") — needs Robert's real emery/gabbro hardware, not done in this sandbox. Builds on #11 (PR #84, state.c's ride state machine) and #58 (PR #81, page.h's persist-key registry / persist_read_data/persist_write_data pattern).
Persist ride state across app restart and crash (#12)
Some checks failed
dev-artifact / build-pbw (push) Failing after 0s
dev-artifact / build-apk (push) Failing after 0s
dev-artifact / publish (push) Has been skipped
fast-lane / host-c-tests (push) Failing after 0s
fast-lane / jvm-tests (push) Failing after 0s
fast-lane / pebble-build (push) Failing after 0s
fast-lane / lint-and-secrets (push) Failing after 0s
fast-lane / meta-declares-required-jobs (push) Failing after 0s
fast-lane / host-c-tests (pull_request) Failing after 0s
fast-lane / jvm-tests (pull_request) Failing after 0s
fast-lane / pebble-build (pull_request) Failing after 0s
fast-lane / lint-and-secrets (pull_request) Failing after 0s
fast-lane / meta-declares-required-jobs (pull_request) Failing after 0s
efdd72ae94
Implements issue #12: an accidental Back press or a crash mid-ride must
not lose the ride. Builds on #11's ride state machine (state.c/state.h)
and #58's persist-key registry (page.h).

watchapp/src/c/state.h / state.c
  - RideStateSnapshot { state, start_epoch, moving_seconds,
    distance_meters }, packed to a fixed 13-byte blob
    (RIDE_STATE_SNAPSHOT_WIRE_SIZE), with the same defensive pack/unpack
    contract page_descriptor_pack()/unpack() already established
    (out-of-range state byte -> reject, *out left untouched).
  - Persisted: state, start_epoch, moving_seconds, distance_meters —
    together "how far into the ride are we". NOT persisted: wallclock_
    seconds (a restart is itself a fresh discontinuity, same as
    ride_state_init() — there's no "current outage" to continue), the
    outbound CMD queue (a queued-but-unsent CMD is stale by the time the
    app restarts and reconnects; the restarted ride reconciles from the
    phone's next RIDE_STATE push instead, same as any other reconnect),
    and connected/stop_pending/protocol-version (live link/UI facts,
    always recomputed at startup). Full reasoning is in state.h's own
    comment ahead of the RideStateSnapshot struct.
  - ride_state_resume_from_snapshot() restores state/start_epoch/
    moving_seconds/distance_meters and re-baselines the tick clock (as if
    ride_state_init() had just run), and now *explicitly* zeroes
    wallclock_seconds/queue_count/stop_pending rather than only assuming
    they're already clean from ride_state_init() — a real gap caught by
    its own new host test (test_resume_from_snapshot_leaves_queue_and_
    wallclock_and_stop_pending_at_init_values), not just a paper
    correctness question.
  - ride_state_should_checkpoint(): write-throttling policy. Due
    immediately after a transition (RIDE_TRANSITION_* other than NONE —
    the moment before a crash matters most), otherwise at most once per
    RIDE_STATE_CHECKPOINT_INTERVAL_S (30 s), and only while RUNNING or
    PAUSED. 30 s bounds a four-hour ride to ~480 writes total (negligible
    against flash endurance) while bounding data loss on a crash to at
    most 30 s of distance/moving-time.
  - ride_state_store_save()/load()/clear(): the only three functions in
    this file touching persist_read_data()/persist_write_data()/
    persist_delete(), under page.h's new PERSIST_KEY_RIDE_STATE — same
    split as page.c's page_store_load()/save(), host-tested the same way
    via tests/stubs/pebble.h's fake persist backing (persist_delete()
    added to the stub for this).

watchapp/src/c/ride_link.c
  - prv_tick_callback() calls ride_state_should_checkpoint(now) once per
    ~1 Hz tick and ride_state_store_save() when it says to — all the
    "when" policy lives in state.c, this file just obeys it.

watchapp/src/c/main.c
  - prv_init() reordered: ride_link_init() (which calls
    ride_state_init()) now runs before carousel_init(), and a new
    prv_offer_resume_if_persisted() runs between them — ride_state_
    store_load() + ride_state_snapshot_is_resumable(), then
    carousel_offer_resume() if there's something to offer. This has to
    happen before carousel_init() pushes the window, so the very first
    render can be the resume prompt instead of the ordinary Ride page.

watchapp/src/c/carousel.h / carousel.c
  - New CONFIRM_RESUME_RIDE case in the existing ConfirmKind/
    prv_confirm_kind()/prv_render() switch — the same prompt mechanism
    CONFIRM_STOP/CONFIRM_EXIT/CONFIRM_PROTOCOL_MISMATCH already use, not
    a new one. "Resume ride?\nSelect=Y\nBack=N" reuses CONFIRM_STOP's
    exact line-length budget rather than a longer, more descriptive
    string that risks wrapping to a clipped fourth line on basalt's
    144 px width.
  - Select -> ride_state_resume_from_snapshot(), logged at INFO so a
    `pebble logs` session after a force-kill/relaunch shows the restored
    values directly. Back -> ride_state_store_clear() (nothing to undo:
    the live state machine was never touched on the discard path), so a
    second restart with no new ride started doesn't offer the same ride
    back up again.

watchapp/tests/ (CMakeLists.txt, stubs/pebble.h, stubs/pebble_stub.c,
test_state.c)
  - persist_delete() added to the fake persist backing.
  - test_state now builds against tests/stubs/pebble.h + pebble_stub.c
    (like test_page already does) since state.c's three store_*()
    functions need it; every other test in the file still exercises pure
    logic with no stub involved.
  - 18 new tests: snapshot pack/unpack round-trip and corrupt-byte
    rejection, is-resumable, resume-from-snapshot (including the queue/
    wallclock/stop_pending gap above), checkpoint throttling (immediate
    on transition, periodic otherwise, only while running/paused, not at
    all while idle/stopped), and store_save/load/clear round-tripping
    through the fake persist backing including wrong-size and corrupt
    contents. 48/48 pass; full suite (fields/page/state/page_render_
    geometry) is 120/120.

Verified:
  - `pebble build` clean on emery, gabbro and basalt (no warnings beyond
    the pre-existing, unrelated RWX-LOAD-segment linker note).
  - All four host test binaries hand-compiled with the exact flags in
    tests/CMakeLists.txt (no cmake binary in this sandbox) and run:
    fields 28/28, page 16/16, state 48/48, page_render_geometry 28/28.
  - Emulator, for real, on emery: installed, started a ride (Select),
    set BT connected so moving time (not wallclock) accumulates,
    confirmed the outbound CMD resend loop in `pebble logs`, then
    SIGKILL'd both the qemu-pebble and pypkjs processes — the harshest
    available force-kill, since a Pebble app has no separate OS process
    of its own to kill independently of the whole emulated watch; only
    the flash-backed persist store survives that. Reinstalling and
    relaunching showed the resume prompt on the very first screen, no
    button pressed yet. Pressing Select logged "Resumed ride: state=1
    start_epoch=1788542315 distance_m=0" — start_epoch matches the
    original CMD_TIME exactly. Repeated the cycle a second time (paused
    ride this time), pressed Back instead: logged "Discarded persisted
    ride snapshot", and a third relaunch showed no prompt at all,
    confirming the discarded snapshot doesn't come back.
  - NOT verified: the physical-watch half of the acceptance criterion
    ("and on the watch"). That needs Robert's real emery/gabbro hardware
    and did not happen in this sandbox — flagged honestly rather than
    claimed, matching every other PR tonight for what still needs
    Robert's devices.

Files: watchapp/src/c/state.c (per the issue), plus state.h, main.c,
carousel.h/c, ride_link.c, and the test harness.

Claude-Session: https://claude.ai/code/session_01DAoXbRmJUf2uxNYBfdAXPt
robert merged commit 977e342af5 into main 2026-09-04 19:23:45 +02:00
Sign in to join this conversation.
No description provided.