Index |  Research ▾  |  Verification ▾  | About
SRCH:FEDF08BF

Impact of High-to-Low Resource Speech Data Ratio on Few-Shot ASR Convergence in Unseen Languages

Submitted: 2 July 2026
Review score: 7.60/10
Verification: L2, Source-grounded claims
Gate status: Verified
Quality tier: DOI grade
Verified claims: 12
DOI: 10.5281/zenodo.21131359

Abstract

Abstract: Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks. However, their applicability is still less explored in low-resource settings. This work investigates the use of Speech LLMs for low-resource Automatic Speech Recognition using the SLAM-ASR framework, where a trainable lightweight projector connects a speech encoder and a LLM. Firstly, we assess training data volume requirements to match Whisper-only performance, re-emphasizing the challenges of limited data. Secondly, we show th

Research Question

How does the ratio of high-resource to low-resource speech data in SLAM-ASR pretraining impact the few-shot word error rate convergence for unseen low-resource languages?

Verification Level

Paper levelL2, Source-grounded claims
Source-grounded claims12
Claim record sourceparsed source sections

Descriptive public verification status only; aggregate claim counts are public, but individual claim records are not exposed here.

Truth-Engine Gate Verdict

StatusVerified
GateGate 2 — Verification (formal proof or sandbox reproduction)
ReasonSealed-sandbox formula repro: Computed 0.009999999999999995 matches expected 0.01 (tolerance=5.0%).
Evaluated2026-07-02T12:53:12.328665+00:00

This record has passed Gate 2: a Lean4 proof source type-checks, or a sealed-sandbox run reproduced the reported results within the stated tolerance. A reproducible artifact (proof source or repro script and results) is attached to this record. VERIFIED requires an attached reproducible artifact (Lean4 proof source, or repro script and results) before this status can be set; it is not derived from review score or claim count.

Quality Tier

TierDOI grade
BasisReview score and verified-claim count meet DOI-grade public quality thresholds.

Descriptive public triage only; this tier does not alter current publication or DOI behavior.

Quality Dimensions

Evidence strength MEDIUM
Citation grounding MEDIUM
Uncertainty disclosure MEDIUM
Reproducibility status HIGH

Automated triage signals derived from public fields; not human peer review or independent validation.

Correction Record

StatusCURRENT
Correction count0
Manifest contractpaper-manifest-v1.1
Correction contractcorrection-record-v1

Public corrections are additive records. Current status does not claim the synthesis is error-free.

Provenance

PublisherAssignee Research
Public provenanceL4, External archival record
Report artifactAvailable
External recordRegistered
Claim lineage12 aggregate source-grounded claims
Review methodAutomated multi-reviewer assessment
Quality guideHow to read scores, claims, manifests, and evidence links
Provenance contractsource-provenance-v1
NoteMachine-generated synthesis of existing literature. Not primary research.