Abstract

On-device assistants run GPU-based LLM inference alongside CPU retrieval on unified-memory systems. Under a saturated local-retrieval workload, four concurrent retrieval workers raise 95th-percentile (p95) decode latency by 60-61% on two M4 systems, whereas prefill latency rises by only 5.7-6.9%. We study LLM phase as an admission signal for independent CPU retrieval under controlled LLM workloads. PHASEGATE calibrates separate concurrency limits for prefill and decode, selecting four and one on our base-M4 configuration. Under a backlogged queue, it achieves 2.0 times the aggregate retrieval throughput of the best tested feasible fixed policy, with both p95 LLM latency metrics within 1.25 times their no-retrieval baselines in all seven held-out runs. A phase-blind control, TimeGate, uses the same two limits on a calibration-derived schedule without observing LLM phase. It achieves similar retrieval throughput but violates the output-token latency limit in every run. M2 and M2 Pro Mac minis reproduce the policy ordering, while output-length sweeps show that the advantage narrows as decode occupies more of each request.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Yum, S., & Kim, S. (2026). PhaseGate: Phase-Aware CPU Retrieval Scheduling for On-Device LLMs on Unified Memory. https://omanscience.com/en/articles/phasegate-phase-aware-cpu-retrieval-scheduling-for-on-device-llms-on-unified-memory

MLA 9

Yum, Seoyoon, and Sehoon Kim. "PhaseGate: Phase-Aware CPU Retrieval Scheduling for On-Device LLMs on Unified Memory." https://omanscience.com/en/articles/phasegate-phase-aware-cpu-retrieval-scheduling-for-on-device-llms-on-unified-memory.

Chicago (author–date)

Yum, Seoyoon, and Sehoon Kim. 2026. "PhaseGate: Phase-Aware CPU Retrieval Scheduling for On-Device LLMs on Unified Memory." https://omanscience.com/en/articles/phasegate-phase-aware-cpu-retrieval-scheduling-for-on-device-llms-on-unified-memory.

Harvard

Yum, S. and Kim, S. (2026) 'PhaseGate: Phase-Aware CPU Retrieval Scheduling for On-Device LLMs on Unified Memory', Available at: https://omanscience.com/en/articles/phasegate-phase-aware-cpu-retrieval-scheduling-for-on-device-llms-on-unified-memory.

Vancouver

Yum S, Kim S. PhaseGate: Phase-Aware CPU Retrieval Scheduling for On-Device LLMs on Unified Memory. https://omanscience.com/en/articles/phasegate-phase-aware-cpu-retrieval-scheduling-for-on-device-llms-on-unified-memory

IEEE

S. Yum, and S. Kim, "PhaseGate: Phase-Aware CPU Retrieval Scheduling for On-Device LLMs on Unified Memory," https://omanscience.com/en/articles/phasegate-phase-aware-cpu-retrieval-scheduling-for-on-device-llms-on-unified-memory.