336 lines
15 KiB
Markdown
336 lines
15 KiB
Markdown
# K720 dispenser issuance contract
|
|
|
|
Hardlink treats the K720 as an unreliable mechanical peripheral with a strict transport
|
|
protocol and imperfect status telemetry. The implementation must keep transport
|
|
validation strict while making mechanical decisions from fresh, bounded observations.
|
|
Diagnostic bytes are useful evidence, but they must not override a valid physical
|
|
position or turn an uncertain observation into a fatal guest-facing error.
|
|
|
|
This document records the dispenser invariants agreed for the clean implementation.
|
|
It is intended to be a design contract: future changes should preserve these rules
|
|
unless production evidence justifies changing them explicitly.
|
|
|
|
## Ownership and concurrency
|
|
|
|
- The dispenser worker is the single owner of serial commands and `deliveryPending`.
|
|
New code must not access the serial port directly from handlers or background
|
|
goroutines.
|
|
- Commands that can move a card are serialized through the worker. There must never
|
|
be concurrent AP/FC7/FC0/RS traffic from separate flows.
|
|
- `deliveryPending` is worker-owned state. It must not be duplicated or mutated by
|
|
HTTP handlers.
|
|
- A real `/issuedoorcard` operation always has priority over maintenance/prestage
|
|
work.
|
|
- Cancellation of a real caller must stop caller-owned work. Internal bounded
|
|
observation timeouts must remain distinguishable from caller cancellation.
|
|
- Do not create overlapping maintenance goroutines or timers per request. Any
|
|
background maintenance must have one clear owner and lifecycle.
|
|
|
|
## Status and observation model
|
|
|
|
AP transport parsing is strict. A response is usable only when its framing and
|
|
contents are valid, including ACK/address/header/length/ETX/BCC/type checks.
|
|
|
|
Malformed, truncated, timed-out or otherwise invalid AP responses are unusable
|
|
observations. They are logged and discarded. They are never converted into a
|
|
fabricated physical state such as `0x30` or `0x38`.
|
|
|
|
Every fresh valid AP observation reclassifies the current mechanical state. No
|
|
physical position is latched across later observations.
|
|
|
|
The fourth AP status byte is the physical position used by the preparation logic:
|
|
|
|
- Encoder confirmed:
|
|
`0x32, 0x33, 0x36, 0x37, 0x3A, 0x3B, 0x3E, 0x3F`.
|
|
A card may be handed to `LockSequence` immediately.
|
|
- Valid but uncertain:
|
|
`0x31, 0x34, 0x35, 0x39, 0x3C, 0x3D`.
|
|
Continue bounded fresh observation; if no stronger state appears, one encoding
|
|
opportunity may still be granted after the uncertainty window.
|
|
- No card on sensors:
|
|
exact `0x30`.
|
|
This may trigger the bounded mechanical recovery sequence.
|
|
- Card well empty:
|
|
exact `0x38`.
|
|
This is the only physical state that becomes `ErrCardWellEmpty`.
|
|
|
|
The first three diagnostic bytes are advisory. Values such as `Preparing card fails`,
|
|
`Dispense card error`, `Card jammed`, or `Command cannot execute` may coexist with a
|
|
usable physical position. They may be logged, but they must not automatically block
|
|
encoding when the fourth byte confirms a usable card position.
|
|
|
|
## Preparing the current card
|
|
|
|
A `/issuedoorcard` request receives at most one `LockSequence` opportunity.
|
|
|
|
Preparation uses bounded fresh observation:
|
|
|
|
- Poll interval: 1 second.
|
|
- After a successful FC7, allow 3 seconds of observation before shake recovery is
|
|
eligible for a persistent exact `0x30`.
|
|
- For uncertain or unusable observations after a successful FC7, use one shared
|
|
4-second uncertainty window. Switching between uncertain and unusable observations
|
|
does not restart that window.
|
|
- A later usable observation always takes precedence:
|
|
encoder confirmed -> handoff;
|
|
exact `0x38` -> empty;
|
|
exact `0x30` -> resume the remaining shake budget and reset uncertainty timing.
|
|
- Total preparation budget: 32 seconds.
|
|
|
|
For persistent exact `0x30`, the request may perform at most three shake recoveries:
|
|
|
|
`RS -> 2 second settle -> FC7`
|
|
|
|
There is no additional plain FC7 retry. Including the initial dispatch, the maximum
|
|
per request is four FC7 commands and three RS commands.
|
|
|
|
If all three shakes are exhausted while the latest usable physical state remains
|
|
exact `0x30`, preparation fails. If the internal preparation deadline is reached
|
|
while the latest state is uncertain or unusable after a successful FC7, the request
|
|
may hand off one encoding opportunity. Exact `0x38` remains empty.
|
|
|
|
FC7 or RS dispatch failures are real transport/command failures. They are not
|
|
converted into status uncertainty and are not retried merely because their dispatch
|
|
failed.
|
|
|
|
## Encoding and delivery
|
|
|
|
`LockSequence` runs exactly once for a prepared card in a single `/issuedoorcard`
|
|
request. The dispenser layer does not implement a second encoding attempt on the same
|
|
physical card.
|
|
|
|
After `LockSequence`, FC0 is dispatched exactly once to move the current card toward
|
|
the guest:
|
|
|
|
- `LockSequence` success -> FC0 once -> successful issuance remains successful.
|
|
- `LockSequence` failure -> FC0 once -> return the original encoding error.
|
|
- FC0 errors are logged only. They do not replace the original encoding result and
|
|
they do not become `ErrCardWellEmpty`.
|
|
|
|
A failed `LockSequence` does not start next-card prestaging.
|
|
|
|
The design intentionally does not add retained-card state, per-card failure counters,
|
|
provider-specific encoder retries, special bad-card recovery, CP/capture recovery, or
|
|
multiple encoding attempts per physical card without production evidence requiring
|
|
them.
|
|
|
|
## Delivery clearance
|
|
|
|
After FC0, the worker marks the previous delivery as pending. A later FC7 must not be
|
|
sent until the previous delivery has been considered clear.
|
|
|
|
Delivery clearance uses fresh AP observations and is owned by the worker.
|
|
|
|
Current bounds:
|
|
|
|
- Minimum clearance wait: 2 seconds.
|
|
- Clearance timeout: 6 seconds.
|
|
- Poll interval: 1 second.
|
|
|
|
A valid physical position of `0x30` or `0x34` may clear `deliveryPending` after the
|
|
minimum wait unless the same fresh observation explicitly indicates active
|
|
preparing/dispensing/capturing movement.
|
|
|
|
Exact `0x38` reports physical empty and does not silently clear the state.
|
|
|
|
If the internal clearance timeout is reached while the caller context is still
|
|
valid, the worker may make the bounded assumption that delivery has cleared unless
|
|
the latest fresh usable observation establishes exact physical empty (`0x38`) or
|
|
explicit active movement. This fallback may therefore occur while the latest
|
|
physical position is `0x33`; `0x33` itself is not positive clearance evidence.
|
|
The fallback is an internal timeout policy, not a reclassification of the observed
|
|
position.
|
|
|
|
If the latest fresh usable observation still explicitly indicates movement, that is
|
|
not treated as unknown. `deliveryPending` remains set and the clearance attempt
|
|
times out.
|
|
|
|
Caller cancellation or deadline always takes precedence over the internal clearance
|
|
fallback. If the caller context expires, return the caller error and do not admit a
|
|
subsequent FC7 from that caller-owned operation.
|
|
|
|
A later unusable observation supersedes older movement evidence; stale movement
|
|
information must not be latched indefinitely.
|
|
|
|
## Next-card prestage and guest UX
|
|
|
|
Preparing the next card is an optimization for the next guest, not part of the
|
|
business success of the current guest's issuance.
|
|
|
|
After a successful `LockSequence` and FC0, Hardlink may make a short best-effort
|
|
attempt to prestage the next card:
|
|
|
|
- wait for worker-owned delivery clearance;
|
|
- if clearance is obtained within the bounded prestage context, send exactly one FC7;
|
|
- do not run readiness polling, RS/shake recovery, or the full current-card
|
|
preparation flow;
|
|
- prestage failure is log-only and must not change the successful `/issuedoorcard`
|
|
result.
|
|
|
|
Do not extend the current guest's screen by 15-20 seconds merely to guarantee that
|
|
the next card reaches the encoder. The guest-facing flow must remain bounded even
|
|
when the dispenser is slow to become ready for prestage.
|
|
|
|
If the short prestage window expires, skip that prestage attempt. A later request or
|
|
maintenance cycle may prepare the next card.
|
|
|
|
## HTTP contract
|
|
|
|
Physical empty and operational failure are deliberately different outcomes.
|
|
|
|
- HTTP 503 is reserved for `ErrCardWellEmpty`, derived only from an exact valid
|
|
physical `0x38`.
|
|
- All other dispenser preparation, transport, command and encoding failures return
|
|
HTTP 502.
|
|
- Normal request/protocol validation keeps its existing 400/405/415 behavior.
|
|
- Diagnostic text such as `Preparing card fails` must never by itself produce 503.
|
|
|
|
Operafyne treats:
|
|
|
|
- 502 as retryable while attempts remain;
|
|
- 503 as terminal `dispenser_failed`;
|
|
- a maximum of three total issue attempts: the initial attempt plus up to two user
|
|
retries.
|
|
|
|
The dispenser layer must preserve this distinction.
|
|
|
|
## Passive status and alerts
|
|
|
|
Ordinary status queries remain strict and passive. They must not reuse tolerant
|
|
issuance semantics to fabricate a status or hide malformed transport.
|
|
|
|
Passive polling captures the current foreground activity generation when the poll is
|
|
queued. The worker skips a passive poll if foreground activity is active when it is
|
|
dispatched or if the captured generation has become obsolete. Passive polling does
|
|
not reset the idle-maintenance clock.
|
|
|
|
Status/diagnostic observations may be useful for logs and support alerts, but alerts
|
|
must not alter physical state classification or issuance success.
|
|
|
|
Idle maintenance observations are intentionally quiet: they do not invoke the normal
|
|
stock-update callback or generate repeated support email such as `Preparing card
|
|
fails`. Maintenance failures are local diagnostic events only.
|
|
|
|
## Idle prestage maintenance
|
|
|
|
Idle prestaging recovers from a short post-FC0 prestage timeout without keeping the
|
|
current guest waiting.
|
|
|
|
The implementation uses the existing dispenser worker as the single maintenance
|
|
owner:
|
|
|
|
- one resettable maintenance timer is owned by the serial-worker loop; there is no
|
|
maintenance goroutine and no timer created per request;
|
|
- `/issuedoorcard` and `/testissuedoorcard` register foreground activity before
|
|
using the dispenser/encoder and release it when the handler finishes;
|
|
- foreground activity is reference-counted so overlapping requests suppress
|
|
maintenance until the last active request finishes;
|
|
- each foreground registration increments an activity generation and invalidates the
|
|
previous idle deadline;
|
|
- when the last foreground request finishes, the next maintenance attempt is
|
|
scheduled one minute later;
|
|
- enabling maintenance at startup schedules the first check one minute later when
|
|
no foreground request is active;
|
|
- a stale timer or stale generation cannot perform maintenance;
|
|
- after a maintenance attempt, the next check is scheduled one minute from completion
|
|
only if the same generation is still authoritative. A later foreground completion
|
|
therefore cannot have its newer deadline overwritten by an older maintenance
|
|
attempt.
|
|
|
|
Foreground registration and maintenance transaction admission use the same activity
|
|
guard. The guard is used only for admission/state bookkeeping and is never held
|
|
across serial I/O, queue waits, sleeps, callbacks, or `LockSequence`.
|
|
|
|
A maintenance attempt is deliberately weaker than foreground preparation:
|
|
|
|
1. create a bounded 5-second maintenance context;
|
|
2. recheck maintenance admission before AP;
|
|
3. obtain one fresh AP using the existing strict transport parser;
|
|
4. classify the fresh status using the common physical classifier;
|
|
5. do nothing for encoder-present, exact `0x38`, unusable, explicit-movement, or
|
|
otherwise ineligible observations;
|
|
6. when `deliveryPending` is set, require fresh qualifying clearance evidence and the
|
|
existing minimum clearance wait before clearing it;
|
|
7. never use the foreground six-second assumed-clearance fallback;
|
|
8. recheck foreground count, activity generation, maintenance enabled state, and
|
|
maintenance context immediately before FC7 admission;
|
|
9. dispatch at most one FC7 and return immediately.
|
|
|
|
If foreground activity begins while an already-admitted maintenance AP is in
|
|
progress, that AP may finish, but the second admission check prevents maintenance
|
|
from sending FC7 afterward. An FC7 that was already admitted and dispatched before
|
|
foreground registration cannot be recalled.
|
|
|
|
Maintenance never performs:
|
|
|
|
- RS or shake recovery;
|
|
- `LockSequence`;
|
|
- `PrepareCurrentCard`;
|
|
- `PrepareNextCard` or `BeginPrepareNextCard`;
|
|
- readiness polling after FC7;
|
|
- the 32-second foreground preparation flow;
|
|
- repeated FC7 attempts;
|
|
- the foreground assumed-clearance fallback.
|
|
|
|
The 5-second maintenance context bounds new admissions and context-aware waits. An
|
|
already-admitted serial read may still finish according to the existing serial read
|
|
timeout, but no subsequent maintenance transaction is admitted after cancellation or
|
|
expiry.
|
|
|
|
`StartMaintenance` and `StopMaintenance` are idempotent lifecycle operations.
|
|
`Client.Close` disables maintenance and stops the timer.
|
|
Stopping maintenance cancels the current maintenance context and waits only for an
|
|
already-admitted maintenance transaction to finish under the existing bounded serial
|
|
behavior.
|
|
|
|
## Automated verification
|
|
|
|
Tests should preserve the behavioural contract rather than only exercise individual
|
|
functions.
|
|
|
|
At minimum, cover:
|
|
|
|
- strict AP framing and malformed/truncated/timeout responses;
|
|
- tolerant issuance observations never fabricating `0x30` or `0x38`;
|
|
- every physical position class and transitions between classes;
|
|
- exact `0x38` as the only `ErrCardWellEmpty` path;
|
|
- persistent `0x30` using at most three `RS -> settle -> FC7` recoveries;
|
|
- no extra FC7 after the shake budget;
|
|
- uncertain/unusable shared timing and later usable-state precedence;
|
|
- caller cancellation versus internal preparation timeout;
|
|
- exactly one `LockSequence` opportunity per request;
|
|
- FC0 exactly once after encoding success or failure;
|
|
- FC0 failure remaining log-only;
|
|
- failed encoding never starting prestage;
|
|
- worker-owned `deliveryPending` and guarded FC7 dispatch;
|
|
- bounded delivery-clearance fallback and explicit-movement timeout;
|
|
- successful issuance remaining successful when next-card prestage fails;
|
|
- HTTP 503 only for exact physical empty and 502 for other issuance failures;
|
|
- Operafyne retry semantics remaining three total attempts.
|
|
|
|
Idle maintenance verification additionally covers:
|
|
|
|
- first attempt occurs one minute after the latest issuance finishes;
|
|
- a new issuance resets that idle interval;
|
|
- overlapping foreground requests suppress maintenance until the last request
|
|
finishes;
|
|
- both `/issuedoorcard` and `/testissuedoorcard` participate in foreground activity
|
|
registration;
|
|
- stale timer events and stale activity generations cannot perform maintenance;
|
|
- foreground registration before maintenance AP prevents AP admission;
|
|
- foreground registration during an admitted AP prevents the later FC7;
|
|
- passive polls are skipped while foreground activity is active or when their
|
|
captured generation is stale;
|
|
- maintenance performs one AP and at most one FC7, with no RS/shake/encoding;
|
|
- encoder-present, exact-empty, movement and unusable observations are no-ops;
|
|
- pending delivery requires fresh normal clearance plus the minimum wait and never
|
|
uses the foreground assumed-clearance fallback;
|
|
- maintenance errors do not produce guest-facing failures, stock callbacks or
|
|
repeated alert email;
|
|
- repeated maintenance start/stop calls are safe and do not accumulate work;
|
|
- cancellation, `StopMaintenance` and `Client.Close` prevent subsequent maintenance
|
|
commands after shutdown admission is revoked.
|
|
|
|
For Hardlink verification, continue excluding the existing test that sends real
|
|
email when running the full suite.
|