الباحثون

Genglin Wang

المنشورات 2

نسخة أولية وصول مفتوح

CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion

Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cr …

نسخة أولية وصول مفتوح

EdgeCraft: Automated Model Crafting for Edge IoT

Genglin Wang, Kaiwei Liu, Liekang Zeng وآخرون · 2026

Machine learning (ML) increasingly powers Internet of Things (IoT) applications at the edge. Yet producing a deployable edge ML artifact for a specific scenario requires navigating a huge search space spanning data representation, model design, training on domain-specific data, and runtime customization. This workflow …

المؤلفون المشاركون