Mobile deep inference is squeezed from three directions at once: the device has little compute or energy to spare, the cloud sees whatever it is sent, and the link between them is unreliable. Split inference, a device prefix plus a cloud remainder, is the usual compromise, but existing runtimes optimise one objective and treat privacy and failure recovery as afterthoughts. AdaSplit handles all three in one control loop. Per inference it selects a sub-layer split in under a millisecond by minimising a weighted latency, energy and privacy objective over 51 precomputed candidates. It obfuscates the transmitted activation with a keyed orthogonal projection and Gaussian noise whose scale is set by a linear-Gaussian bound on reconstruction error (Proposition 1), which we treat as a calibration knob rather than a security guarantee. When the link drops, the device resends the buffered, already obfuscated activation instead of recomputing the prefix. Across three bandwidth regimes on an emulated CIFAR-10 and ResNet-56 testbed, against five baselines and two ablations, AdaSplit matches the median latency of the best split baseline (27.1 to 29.6 ms), raises the reconstruction error of a key-less inversion adversary by 83 times (PSNR 42.25 to 20.83 dB), stays within 1.1 points of the 94.2 percent cloud-only ceiling where uncalibrated noise collapses to 73 percent, and eliminates the 28 to 47 percent of device work that a drop wastes. Recovery buys that robustness with tail latency, which we measure rather than omit. Latency and energy are emulation proxies, so this is a proof of concept for the three mechanisms rather than a deployment claim.
发表评论