AdaSplit: Latency, Privacy, and Fault Tolerance Without Trading One for Another in Edge-Cloud Inference
编号:48 访问权限:仅限参会人 更新:2026-10-04 23:25:29 浏览:9次 Online

报告开始:2026年10月13日 14:45(Asia/Ho_Chi_Minh)

报告时间:15min

所在会场:[S5] Track 5: Emerging Trends of AI/ML [S5-6] Track 5: Emerging Trends of AI/ML

暂无文件

摘要
Mobile deep inference is squeezed from three directions at once: the device has little compute or energy to spare, the cloud sees whatever it is sent, and the link between them is unreliable. Split inference, a device prefix plus a cloud remainder, is the usual compromise, but existing runtimes optimise one objective and treat privacy and failure recovery as afterthoughts. AdaSplit handles all three in one control loop. Per inference it selects a sub-layer split in under a millisecond by minimising a weighted latency, energy and privacy objective over 51 precomputed candidates. It obfuscates the transmitted activation with a keyed orthogonal projection and Gaussian noise whose scale is set by a linear-Gaussian bound on reconstruction error (Proposition 1), which we treat as a calibration knob rather than a security guarantee. When the link drops, the device resends the buffered, already obfuscated activation instead of recomputing the prefix. Across three bandwidth regimes on an emulated CIFAR-10 and ResNet-56 testbed, against five baselines and two ablations, AdaSplit matches the median latency of the best split baseline (27.1 to 29.6 ms), raises the reconstruction error of a key-less inversion adversary by 83 times (PSNR 42.25 to 20.83 dB), stays within 1.1 points of the 94.2 percent cloud-only ceiling where uncalibrated noise collapses to 73 percent, and eliminates the 28 to 47 percent of device work that a drop wastes. Recovery buys that robustness with tail latency, which we measure rather than omit. Latency and energy are emulation proxies, so this is a proof of concept for the three mechanisms rather than a deployment claim.
关键词
Terms—Adaptive model partitioning, edge-cloud collaborative inference, fault tolerance, feature inversion attacks, mobile deep learning, privacy-preserving inference, split computing.
报告人
Goutham Reddya
Staff Engering Independent Research

稿件作者
Goutham Reddya Independent Research
发表评论
验证码 看不清楚,更换一张
全部评论
重要日期
  • 会议日期

    10月11日

    2026

    至

    10月14日

    2026

  • 12月30日 2025

    报告提交截止日期

  • 09月28日 2026

    提前注册日期

  • 10月10日 2026

    初稿截稿日期

  • 10月14日 2026

    注册截止日期

主办单位
United Societies of Science
承办单位
Posts and Telecommunications Institute of Technology
协办单位
IEEE Section
IEEE Vietnam Section
移动端
在手机上打开
小程序
打开微信小程序
客服
扫码或点此咨询