To what extent does expert-level quantization (e.g., INT8 vs FP16) combined with CPU-GPU expert caching reduce

Assignee Research

SRCH:3E885544

To what extent does expert-level quantization (e.g., INT8 vs FP16) combined with CPU-GPU expert caching reduce

Submitted: 28 May 2026
Review score: 4.83/10
Verification: L1, Literature synthesis
Quality tier: Quarantine candidate

PDF BibTeX RIS Manifest Corrections

Abstract

Abstract: Among parallel decoding paradigms, diffusion large language models (dLLMs) have emerged as a promising candidate that balances generation quality and throughput. However, their integration with Mixture-of-Experts (MoE) architectures is constrained by an expert explosion: as the number of tokens generated in parallel increases, the number of distinct experts activated grows nearly linearly. This results in substantial memory traffic that pushes inference into a memory-bound regime, negating the efficiency gains of both MoE and parallel decoding. To address this challenge, we propose Dynamic Exp

Research Question

To what extent does expert-level quantization (e.g., INT8 vs FP16) combined with CPU-GPU expert caching reduce memory transfer overhead in MoE diffusion LLMs, and what is the trade-off in generation accuracy on the GSM8K and HumanEval benchmarks?

Verification Level

Paper level	L1, Literature synthesis
Source-grounded claims	0
Claim record source	not publicly specified

Descriptive public verification status only; aggregate claim counts are public, but individual claim records are not exposed here.

Quality Tier

Tier	Quarantine candidate
Basis	Review score is below 5.0; source-level inspection is required before relying on the synthesis.

Descriptive public triage only; this tier does not alter current publication or DOI behavior.

Quality Dimensions

Evidence strength	LOW
Uncertainty disclosure	MEDIUM
Reproducibility status	MEDIUM

Automated triage signals derived from public fields; not human peer review or independent validation.

Correction Record

Status	CURRENT
Correction count	0
Manifest contract	paper-manifest-v1.1
Correction contract	correction-record-v1

Public corrections are additive records. Current status does not claim the synthesis is error-free.

Provenance

Publisher	Assignee Research
Public provenance	L2, Public artifact record
Report artifact	Available
External record	Not registered
Claim lineage	0 aggregate source-grounded claims
Review method	Automated multi-reviewer assessment
Quality guide	How to read scores, claims, manifests, and evidence links
Provenance contract	source-provenance-v1
Note	Machine-generated synthesis of existing literature. Not primary research.