Efficient LLM Inference
KV-cache reuse, semantic cache distillation, state transfer, selective recomputation, and bandwidth-aware serving.
PH.D. STUDENT — INST. OF AI & FUTURE NETWORKS, BNU
~$
I work on efficient LLM systems, semantic state transfer, multi-candidate reasoning, and security-oriented machine learning — building AI that is fast to serve and safe to trust.
01 /
From KV-cache level systems optimization to model-level security analysis — one goal: AI that is fast, cheap, and trustworthy.
KV-cache reuse, semantic cache distillation, state transfer, selective recomputation, and bandwidth-aware serving.
Candidate construction, fixed-verifier reranking, coverage-conversion gaps, and answer-mode compatibility.
Backdoor attacks, frequency-domain robustness, model behavior under perturbation, and security evaluation.
Vision pipelines for smart livestock systems, identity-related signals, weight estimation, and agricultural AI.
02 /
Recent work on cache distillation, multi-candidate reasoning, and frequency-domain backdoors.
Proc. 43rd International Conference on Machine Learning
Under review
03 /
Systems-oriented research with applied ML depth.
Designed REUSE and PATCH mechanisms for state transfer between shared-architecture, weight-mismatched models.
Studied stealthy backdoor mechanisms, trigger design, and robustness-oriented model analysis.
Built applied pipelines for cattle image acquisition, weight estimation, visual signal analysis, and farm decision support.
04 /