Selected as one of four contributed talks (oral presentation) at the workshop.
A compromised LLM inference server can leak model weights by encoding payload bits in otherwise plausible token choices. A replay of the same prompt in a trusted server can expose such deviations, but benign numerical nondeterminism also causes token mismatches, so patient attackers can hide within normal variation unless evidence is combined across responses.
This work introduces a prompt-level e-process that calibrates whole-response mismatch events on trusted benign traffic and accumulates evidence sequentially while, under a calibration-transfer assumption, controlling the probability of any false alarm over an unbounded monitoring horizon. The construction combines nested margin events, Clopper-Pearson calibration of their benign rates, and a betting-style sequential update whose validity follows from Ville’s inequality.
The monitor is evaluated on four models against a seed-blind attack and a stronger seed-aware attack that hides payload bits only in near-ties to remain stealthy, analyzing the channel capacity versus detectability trade-off. Compared with a hard per-token alarm, the e-process combines weak evidence across responses while providing explicit anytime false-alarm control.