Index |  Research ▾  |  Verification ▾  | About
SRCH:E9E7D370

DeepSeek-V3 Parameter Scaling and Accuracy Variance on GPQA Diamond Under Distribution Shifts

Submitted: 30 May 2026
Review score: 3.17/10
Verification: L1, Literature synthesis
Gate status: Unverified
Quality tier: Quarantine candidate

Abstract

Abstract: This report synthesises findings from 13 peer-reviewed papers addressing the following research question: How does increasing parameter count from 7B to 33B in DeepSeek-V3 affect accuracy variance on GPQA Diamond under synthetic distribution shifts. In electronic trading markets, limit order books (LOBs) provide information about pending buy/sell orders at various price levels for a given security. Recently, there has been a growing interest in using LOB data for resolving downstream machine learning tasks (e.g.. 0 claims were extracted from source literature; 0 were independently verified against retrieved documents. An automated multi-reviewer quality assessment produced a score of 3.2/10. This report is a machine-generated literature synthesis and does not constitute original research.

Research Question

How does increasing parameter count from 7B to 33B in DeepSeek-V3 affect accuracy variance on GPQA Diamond under synthetic distribution shifts?

Verification Level

Paper levelL1, Literature synthesis
Source-grounded claims0
Claim record sourceparsed source sections

Descriptive public verification status only; aggregate claim counts are public, but individual claim records are not exposed here.

Truth-Engine Gate Verdict

StatusUnverified
GateGate 2 — Verification (formal proof or sandbox reproduction)
ReasonPublished before the Gate 2 verification pipeline was activated (2026-06-10). No formal proof or sandbox reproduction has been attempted for this record.
Evaluated2026-06-10T06:30:49+00:00

This record has not completed Gate 2 of the verification pipeline (a type-checked Lean4 proof for mathematical claims, or a sealed-sandbox reproduction for empirical claims). It is a literature synthesis only. VERIFIED requires an attached reproducible artifact (Lean4 proof source, or repro script and results) before this status can be set; it is not derived from review score or claim count.

Quality Tier

TierQuarantine candidate
BasisReview score is below 5.0; source-level inspection is required before relying on the synthesis.

Descriptive public triage only; this tier does not alter current publication or DOI behavior.

Quality Dimensions

Evidence strength LOW
Uncertainty disclosure MEDIUM
Reproducibility status MEDIUM

Automated triage signals derived from public fields; not human peer review or independent validation.

Correction Record

StatusCURRENT
Correction count0
Manifest contractpaper-manifest-v1.1
Correction contractcorrection-record-v1

Public corrections are additive records. Current status does not claim the synthesis is error-free.

Provenance

PublisherAssignee Research
Public provenanceL2, Public artifact record
Report artifactAvailable
External recordNot registered
Claim lineage0 aggregate source-grounded claims
Review methodAutomated multi-reviewer assessment
Quality guideHow to read scores, claims, manifests, and evidence links
Provenance contractsource-provenance-v1
NoteMachine-generated synthesis of existing literature. Not primary research.