Public Falsification Record
Gate 2 / Gate 3 Pipeline Falsifications (1052)
Claims VERIFIED at Gate 2 (sealed-sandbox repro) and subsequently falsified by the Gate 3 adversarial red-team (three independent LLM attackers, inverted scoring). A claim SURVIVES only if all three attackers fail to find a fatal flaw (avg attack score < 3.5; no individual score ≥ 5.0).
| Task ID | Gate | Claim type | Goal / Claim | Avg attack | Killed (UTC) |
|---|---|---|---|---|---|
| 2fc25024-9375-43… | Gate 3 | formula_repro |
How does fine-tuning on intermediate tasks with varying levels of linguistic diversity (monolingual vs. multilingual intermediate tasks) aff…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-07-26 02:17 |
| 3b0437e6-24b1-4b… | Gate 3 | formula_repro |
How does the performance of multilingual intermediate-task training compare to monolingual English intermediate-task training when evaluated…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-26 02:17 |
| 863bd782-4975-46… | Gate 3 | formula_repro |
How does the performance gain from English intermediate-task training vary across different language families in XTREME-R tasks when compari…
COUNTEREXAMPLE HUNTER: 7.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-26 02:17 |
| 2e405f49-aad8-4a… | Gate 3 | formula_repro |
How does scaling the size of intermediate-task datasets impact the zero-shot cross-lingual transfer performance on XTREME-R, comparing accur…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.4/10 | 2026-07-26 02:16 |
| de83404c-3d90-4c… | Gate 3 | formula_repro |
How does increasing XLM-R model size from base to large affect zero-shot accuracy on XTREME-R natural language inference tasks following Eng…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-07-26 02:16 |
| 8f509ea4-80c2-46… | Gate 3 | formula_repro |
What is the effect of varying the acoustic diversity in pre-training data on the accuracy of self-supervised speech models, as measured by W…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.5/10 | 2026-07-26 02:16 |
| 6c1607c2-6831-46… | Gate 3 | formula_repro |
What is the impact of domain shift in intermediate tasks (e.g., news vs. social media tasks) on zero-shot cross-lingual transfer performance…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
4.7/10 | 2026-07-26 02:16 |
| c41766d6-b154-45… | Gate 3 | formula_repro |
What is the impact of intermediate-task dataset size scaling on zero-shot cross-lingual transfer performance for multilingual models on the …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-07-25 20:16 |
| a0889b90-bf6c-43… | Gate 3 | formula_repro |
How does the scale of intermediate-task data impact zero-shot cross-lingual transfer performance on XTREME-R when using intermediate-task tr…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.5/10 | 2026-07-25 20:16 |
| 31d78180-97b5-44… | Gate 3 | formula_repro |
How does the scaling of model size (e.g., 7B, 13B, 30B parameters) affect the effectiveness of English intermediate-task training for zero-s…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-07-25 14:16 |
| eb0b73f5-f8f5-4f… | Gate 3 | formula_repro |
Does combining intermediate-task training with parameter-efficient fine-tuning (PEFT) methods improve zero-shot cross-lingual transfer perfo…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-25 14:15 |
| c6dc89a8-8c5f-48… | Gate 3 | formula_repro |
What is the impact of data augmentation with synthetic parallel corpora on the robustness of zero-shot cross-lingual transfer for multimodal…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-07-25 14:15 |
| 745e46d6-ab64-4c… | Gate 3 | formula_repro |
Does the performance gain from intermediate-task training on XTREME-R benchmark tasks scale with the model size when evaluated using XTREME-…
COUNTEREXAMPLE HUNTER: 8.0/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 2.5/10
|
5.9/10 | 2026-07-25 14:15 |
| 7af7498e-7263-40… | Gate 3 | formula_repro |
Does intermediate-task fine-tuning on English improve reasoning capabilities in non-English XTREME tasks, as measured by accuracy on XTREME-…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 6.5/10
|
5.7/10 | 2026-07-25 14:15 |
| b21edad3-06c9-4c… | Gate 3 | formula_repro |
Does scaling the number of intermediate language-understanding tasks improve zero-shot cross-lingual transfer performance on XTREME, and at …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-25 14:14 |
| c721a404-0c69-44… | Gate 3 | formula_repro |
How does the choice of intermediate task (e.g., NLI, QA, or summarization) impact zero-shot cross-lingual transfer on the XTREME-R benchmark…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.8/10 | 2026-07-25 08:14 |
| 7390df44-b208-4b… | Gate 3 | formula_repro |
How does scaling the size of the intermediate English task dataset affect the zero-shot cross-lingual transfer performance of multilingual m…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.3/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-07-25 08:14 |
| 81e4b6d9-ab3f-41… | Gate 3 | formula_repro |
How does cross-task consistency in intermediate-task selection (e.g., using only question-answering vs. mixed tasks) affect zero-shot cross-…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-25 08:14 |
| 6f7fdc19-6d5a-47… | Gate 3 | formula_repro |
Does increasing the number of intermediate language-understanding tasks beyond nine in sequential training improve zero-shot cross-lingual t…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.5/10 | 2026-07-25 08:14 |
| 89bf9632-97ce-47… | Gate 3 | formula_repro |
What is the effect of intermediate-task diversity (e.g., combining multiple tasks) on zero-shot cross-lingual transfer performance for non-E…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-25 08:14 |
| fd707850-0000-47… | Gate 3 | formula_repro |
Does scaling the number of intermediate tasks in English before final task fine-tuning impact zero-shot cross-lingual performance on XTREME-…
COUNTEREXAMPLE HUNTER: 6.2/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.1/10 | 2026-07-25 08:14 |
| ae6250be-afae-4d… | Gate 3 | formula_repro |
What is the impact of varying the size of the pretrained model (e.g., base vs. large vs. huge models) on the effectiveness of English interm…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 3.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-25 02:14 |
| e8ca2d82-5f0c-43… | Gate 3 | formula_repro |
How does the computational efficiency (FLOPs/inference time) of English intermediate-task training compare to multimodal intermediate-task t…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
5.3/10 | 2026-07-25 02:14 |
| 541eeacb-0542-4a… | Gate 3 | formula_repro |
Does adapting intermediate-task training to multilingual instruction datasets (e.g., M3IA) outperform English-only training for zero-shot XT…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.0/10
|
5.7/10 | 2026-07-25 02:14 |
| 2c4fc4a6-9e60-44… | Gate 3 | formula_repro |
Does scaling the size of the pre-trained multilingual model affect the performance of intermediate-task training on zero-shot cross-lingual …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-25 02:14 |
| 69752514-f294-4e… | Gate 3 | formula_repro |
How does the choice of English intermediate language understanding task influence the effectiveness of zero-shot cross-lingual transfer on X…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.2/10
|
7.4/10 | 2026-07-24 20:10 |
| 38f21cdd-c051-47… | Gate 3 | formula_repro |
How does the intermediate-task training on English affect the zero-shot performance of multimodal models like M-UNITER on the XTREME-M bench…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-24 20:10 |
| c005cb1b-9cc9-4e… | Gate 3 | formula_repro |
How does incorporating intermediate tasks from multiple languages (beyond English) affect the zero-shot cross-lingual transfer performance a…
COUNTEREXAMPLE HUNTER: 7.0/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.2/10
|
5.9/10 | 2026-07-24 20:10 |
| 5b383f1e-7747-43… | Gate 3 | formula_repro |
How does the choice of intermediate-task complexity (e.g., XNLI vs. MLQA) affect zero-shot cross-lingual transfer performance on XTREME-R, a…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 6.5/10
|
5.7/10 | 2026-07-24 20:10 |
| 8fbff354-f00a-40… | Gate 3 | formula_repro |
Does intermediate-task training on English code-related tasks (e.g., CodeXGLUE) improve zero-shot cross-lingual transfer for programming-rel…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-24 20:10 |
| 377eccb9-4031-46… | Gate 3 | formula_repro |
How does the zero-shot cross-lingual performance of XLM-R Base change when using intermediate-task fine-tuning on code-related tasks (e.g., …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-07-24 20:10 |
| 76089056-2e35-4e… | Gate 3 | formula_repro |
Does replacing English intermediate-task training with multilingual tasks (e.g., mXGLUE) improve exact match scores on XTREME's non-English …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.3/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-24 20:10 |
| 1f04ba9c-8273-43… | Gate 3 | formula_repro |
Does intermediate-task training with code generation benchmarks (e.g., HumanEval, MBPP) improve zero-shot cross-lingual transfer performance…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-07-24 20:10 |
| 3e4bd9d3-cc90-4e… | Gate 3 | formula_repro |
What is the effectiveness of multitask intermediate training on cross-lingual transfer, as measured by average accuracy on XTREME tasks when…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.4/10 | 2026-07-24 14:10 |
| 92950208-ce92-4a… | Gate 3 | formula_repro |
What is the impact of model size (e.g., 7B vs. 13B parameters) on zero-shot cross-lingual transfer performance for English-trained intermedi…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-07-24 14:10 |
| d845176f-ec65-4c… | Gate 3 | formula_repro |
How does the choice of intermediate task diversity (e.g., natural language inference, question answering, sentiment analysis) impact the zer…
COUNTEREXAMPLE HUNTER: 7.3/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
4.6/10 | 2026-07-24 14:10 |
| c7e8092e-0b96-43… | Gate 3 | formula_repro |
How does the choice of intermediate task difficulty (e.g., complexity, data size) affect the zero-shot cross-lingual transfer performance on…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
4.9/10 | 2026-07-24 14:10 |
| acf7e2c4-6450-47… | Gate 3 | formula_repro |
Does combining multiple English intermediate tasks in a multi-task learning setup improve zero-shot transfer on XTREME-R compared to single-…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-07-24 14:10 |
| 40ce3b98-ddc0-4e… | Gate 3 | formula_repro |
How does the choice of intermediate task difficulty (e.g., easy vs. hard classification tasks) affect the zero-shot cross-lingual transfer p…
COUNTEREXAMPLE HUNTER: 5.0/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.6/10 | 2026-07-24 14:10 |
| a8f3ec61-1385-43… | Gate 3 | formula_repro |
How does combining intermediate language-understanding tasks with code generation tasks (e.g., using HumanEval) influence zero-shot cross-li…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.3/10 · REPLICATION ATTACKER: 6.5/10
|
7.1/10 | 2026-07-24 14:10 |
| d158e633-ac02-4e… | Gate 3 | formula_repro |
How does intermediate-task training on logical reasoning benchmarks (e.g., LSAT, DROP) compare to standard language understanding tasks (e.g…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.3/10 | 2026-07-24 14:10 |
| e930b0ff-4938-4f… | Gate 3 | formula_repro |
What is the impact of using domain-specific intermediate tasks versus general language understanding tasks on the zero-shot cross-lingual tr…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-24 14:10 |
| 55ec26c5-22c9-48… | Gate 3 | formula_repro |
How does the performance of XLM-R-Base on zero-shot cross-lingual classification tasks vary when intermediate-task training is applied to do…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-24 08:09 |
| b3971733-3d1b-46… | Gate 3 | formula_repro |
Do multimodal intermediate tasks (e.g., Flickr30k, MMMU) improve zero-shot cross-lingual transfer on the XTREME benchmark more effectively t…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-07-24 08:09 |
| 47b81473-2258-47… | Gate 3 | formula_repro |
How does the trade-off between model size and inference efficiency impact zero-shot cross-lingual performance on the XTREME-R benchmark when…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-07-24 08:09 |
| a919f94f-d4d0-48… | Gate 3 | formula_repro |
Does fine-tuning on intermediate multimodal reasoning tasks (e.g., VQA, visual entailment) improve zero-shot cross-lingual performance on XT…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.3/10 | 2026-07-24 08:09 |
| 7fbe0363-9852-46… | Gate 3 | formula_repro |
Can intermediate-task training on code generation benchmarks (e.g., HumanEval, MBPP) enhance zero-shot cross-lingual reasoning capabilities …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-24 08:09 |
| ef11a965-5f57-46… | Gate 3 | formula_repro |
How does intermediate-task training on multimodal benchmarks (e.g., VQA, CLIP) compare to text-only intermediate tasks in improving zero-sho…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.3/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-24 08:09 |
| 2cabee14-e2d0-48… | Gate 3 | formula_repro |
What is the effect of incorporating multilingual intermediate tasks alongside English tasks on zero-shot cross-lingual transfer performance …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-24 08:09 |
| 8532bd40-9c0f-4e… | Gate 3 | formula_repro |
Does the use of multimodal intermediate tasks (e.g., combining vision-language tasks) further improve zero-shot cross-lingual transfer perfo…
COUNTEREXAMPLE HUNTER: 8.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.3/10 | 2026-07-24 08:09 |
| d299ca72-1510-4f… | Gate 3 | formula_repro |
How does intermediate-task training on non-English tasks specifically affect the performance of multilingual models on XTREME-R benchmarks c…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-24 08:09 |
| b3ce2356-2e48-4d… | Gate 3 | formula_repro |
Does English intermediate-task training improve zero-shot transfer accuracy on XTREME-R for code generation and reasoning tasks compared to …
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 3.5/10
|
6.1/10 | 2026-07-24 03:38 |
| f98f9b6e-cf25-45… | Gate 3 | formula_repro |
Do multimodal models exhibit similar zero-shot cross-lingual transfer benefits from English intermediate-task training as text-only models w…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-07-24 03:38 |
| 582f441f-f813-41… | Gate 3 | formula_repro |
How does domain-specific intermediate-task training (e.g., science vs. social studies tasks) influence zero-shot cross-lingual performance o…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
4.7/10 | 2026-07-24 03:38 |
| 9e4faa5c-06c4-41… | Gate 3 | formula_repro |
What is the effect of different intermediate language understanding task types (e.g., natural language inference, question answering) on the…
COUNTEREXAMPLE HUNTER: 7.3/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-07-24 03:38 |
| 9ba21463-ab51-45… | Gate 3 | formula_repro |
How does intermediate-task training on code generation datasets compare to natural language inference (NLI) tasks in improving zero-shot per…
COUNTEREXAMPLE HUNTER: 7.0/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.5/10 | 2026-07-24 03:38 |
| c517f880-9e0a-46… | Gate 3 | formula_repro |
How does intermediate-task training on multilingual tasks compare to English-only intermediate tasks in terms of zero-shot cross-lingual tra…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-24 03:38 |
| 51202258-2335-49… | Gate 3 | formula_repro |
What is the impact of using multilingual intermediate tasks (instead of English-only) on zero-shot cross-lingual transfer performance, as me…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-24 03:37 |
| 50b20765-3abe-48… | Gate 3 | formula_repro |
How does the combination of intermediate-task training with instruction fine-tuning affect zero-shot performance on XTREME-R compared to eit…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-24 03:37 |
| 81dd6df2-6006-42… | Gate 3 | formula_repro |
How does varying the number of intermediate language understanding tasks during fine-tuning affect the zero-shot cross-lingual transfer perf…
COUNTEREXAMPLE HUNTER: 7.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
4.8/10 | 2026-07-23 21:37 |
| 1dc0672d-5606-47… | Gate 3 | formula_repro |
What is the impact of task diversity (e.g., question answering vs. natural language inference) in intermediate-task training on cross-lingua…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.4/10 | 2026-07-23 21:37 |
| cf7f97bb-c4e7-4e… | Gate 3 | formula_repro |
Does the use of distilled intermediate-task models (e.g., task-specific student models) improve inference efficiency for zero-shot cross-lin…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.4/10 | 2026-07-23 21:37 |
| 81b09b94-de94-45… | Gate 3 | formula_repro |
Does intermediate-task training on English-only datasets improve the reasoning capabilities of multilingual models like XLM-R on non-English…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-23 21:37 |
| 7f8d1d53-2df8-4a… | Gate 3 | formula_repro |
Does scaling the number of intermediate language-understanding tasks from nine to a larger set (e.g., 20+) further improve zero-shot cross-l…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
4.7/10 | 2026-07-23 09:29 |
| 0f3f465f-cbb7-4d… | Gate 3 | formula_repro |
How does scaling the size of the language model affect the zero-shot cross-lingual transfer performance on XTREME-R when using intermediate-…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 0.0/10
|
5.2/10 | 2026-07-23 03:28 |
| 86204aaf-27f3-47… | Gate 3 | formula_repro |
How does intermediate-task training on multilingual code generation datasets impact zero-shot cross-lingual transfer accuracy on the XTREME-…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.4/10 | 2026-07-23 03:28 |
| f745deb8-b652-41… | Gate 3 | formula_repro |
Does incrementally scaling the size of English legal corpora used for intermediate-task training improve zero-shot cross-lingual transfer ac…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
4.9/10 | 2026-07-22 21:27 |
| dcdd45e6-90b9-45… | Gate 3 | formula_repro |
To what extent does the size of the intermediate-task training dataset influence the zero-shot cross-lingual transfer performance of mT5 on …
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 6.5/10
|
7.0/10 | 2026-07-22 15:26 |
| 28544760-975b-4a… | Gate 3 | formula_repro |
How does the scaling of intermediate-task training (e.g., increasing the number of intermediate tasks or model size) affect zero-shot cross-…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 9.2/10
|
7.0/10 | 2026-07-22 15:24 |
| 718bcd09-f12c-40… | Gate 3 | formula_repro |
How does the choice of intermediate-task difficulty (measured by English benchmark accuracy) impact the zero-shot cross-lingual transfer per…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.4/10 | 2026-07-22 15:06 |
| 6bbcbae9-49a6-47… | Gate 3 | formula_repro |
How does the impact of intermediate-task transfer on zero-shot cross-lingual performance compare when using distilled models (e.g., mT5-Smal…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.4/10 | 2026-07-22 15:05 |
| 81deb642-2ee8-46… | Gate 3 | formula_repro |
What is the effect of combining multiple intermediate language-understanding tasks (e.g., NLI, QA, sentiment analysis) on zero-shot cross-li…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.2/10
|
7.6/10 | 2026-07-22 15:04 |
| 5eef4986-e198-49… | Gate 3 | formula_repro |
How does the alignment of intermediate-task training objectives with target task domains (e.g., reasoning vs. code generation) influence the…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 2.5/10
|
5.8/10 | 2026-07-22 15:03 |
| 3b490ff7-2c36-46… | Gate 3 | formula_repro |
Does scaling the size of the intermediate-task training data improve zero-shot cross-lingual transfer performance on the XTREME-R benchmark,…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.2/10 | 2026-07-22 15:02 |
| 563a8c85-d7fc-45… | Gate 2 | unknown | How does the scaling of XLM-R model size (base vs. large) affect the F1 score stability of euphemism detection when using direct cross-lingu… | - | 2026-07-22 14:47 |
| e73e324e-cf29-42… | Gate 3 | formula_repro |
How does the scaling of model size from 1B to 70B parameters impact the zero-shot cross-lingual transfer performance on XTREME-R when using …
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 9.5/10
|
6.5/10 | 2026-07-22 08:58 |
| c3b691ee-545f-46… | Gate 3 | formula_repro |
How does intermediate-task fine-tuning on high-resource English NLU datasets impact zero-shot accuracy on the XTREME-R benchmark compared to…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.5/10 | 2026-07-22 08:57 |
| 84a26c6d-9bcd-4e… | Gate 3 | formula_repro |
How does intermediate-task training on non-English language understanding tasks (e.g., TyDi QA, mT5) affect zero-shot cross-lingual transfer…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.8/10 · REPLICATION ATTACKER: 9.2/10
|
8.2/10 | 2026-07-22 08:53 |
| 52a5460e-4e79-4d… | Gate 3 | formula_repro |
Can the effectiveness of English intermediate-task training for zero-shot cross-lingual transfer be scaled with larger models (e.g., 10B+ pa…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-22 08:49 |
| a0e0e9be-abef-4c… | Gate 3 | formula_repro |
What is the impact of using intermediate tasks from different domains (e.g., reasoning, code generation, multimodal) on the zero-shot cross-…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-22 08:48 |
| c8b347fc-7e70-45… | Gate 3 | formula_repro |
Do models pretrained with English intermediate tasks and then fine-tuned on the XTREME benchmark achieve higher zero-shot cross-lingual accu…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-22 08:47 |
| b6c0f8f0-26ef-47… | Gate 3 | formula_repro |
Does the performance improvement from English intermediate-task training on XNLI transfer to other zero-shot cross-lingual benchmarks like X…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.6/10 | 2026-07-22 08:46 |
| 6292181e-ac14-4a… | Gate 3 | formula_repro |
What is the impact of intermediate-task training on few-shot cross-lingual performance (e.g., 3-shot or 5-shot) across languages in XTREME, …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.5/10 | 2026-07-22 08:44 |
| dae24db4-7f60-44… | Gate 3 | formula_repro |
How does the zero-shot transfer accuracy on XTREME-R compare when using intermediate-task training on multimodal datasets like VQA versus te…
COUNTEREXAMPLE HUNTER: 7.3/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.6/10 | 2026-07-22 01:27 |
| 1e075903-8a8f-43… | Gate 3 | formula_repro |
Does fine-tuning XLM-R Base on intermediate English-only tasks (e.g., GLUE, SuperGLUE) before target task fine-tuning improve zero-shot cros…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-07-22 01:27 |
| c8c5b733-8cd7-4b… | Gate 3 | formula_repro |
Does the inclusion of multimodal intermediate tasks (e.g., image-captioning or visual question answering) in English improve zero-shot perfo…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-07-22 01:27 |
| b20b70e6-0b9c-4c… | Gate 3 | formula_repro |
How does the performance of intermediate-task training on multimodal datasets compare to monolingual English baselines when evaluated on the…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 3.5/10
|
5.8/10 | 2026-07-22 01:27 |
| b5e1ef68-0b08-47… | Gate 3 | formula_repro |
Does scaling the size of the pretrained model improve the robustness of zero-shot cross-lingual transfer performance in the XTREME-R benchma…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-07-22 01:27 |
| d91ebd41-b102-41… | Gate 3 | formula_repro |
How does the inclusion of multimodal intermediate tasks (e.g., image-text alignment) impact zero-shot cross-lingual transfer performance on …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.4/10 | 2026-07-22 01:27 |
| 186d5ba8-62e4-4a… | Gate 3 | formula_repro |
How does the size and diversity of intermediate-task datasets affect the zero-shot cross-lingual accuracy on XTREME-R natural language infer…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-22 01:27 |
| d98bb67c-f10f-42… | Gate 3 | formula_repro |
How does intermediate-task training on code generation datasets affect zero-shot cross-lingual transfer performance on XTREME-R tasks compar…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.4/10 | 2026-07-21 19:26 |
| 8912e919-310b-4a… | Gate 3 | formula_repro |
What is the impact of English intermediate-task training on the semantic reasoning capabilities of multilingual models when evaluated on non…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.3/10 · REPLICATION ATTACKER: 7.2/10
|
7.0/10 | 2026-07-21 19:26 |
| 5e3159e9-9224-42… | Gate 3 | formula_repro |
What is the impact of varying the size of the intermediate-task training dataset on the zero-shot cross-lingual transfer performance of mode…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.4/10 | 2026-07-21 19:26 |
| 6bd9a328-9a00-47… | Gate 3 | formula_repro |
Does scaling the size of pre-trained language models (e.g., Llama 3, PaLM 2) improve the zero-shot cross-lingual transfer performance of int…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.2/10
|
6.0/10 | 2026-07-21 19:23 |
| f541dc44-db91-4c… | Gate 3 | formula_repro |
Does scaling the number of English intermediate tasks for mT5 lead to improved zero-shot cross-lingual performance on XTREME-R, as measured …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-07-21 13:07 |
| 07c7291b-d10e-49… | Gate 3 | formula_repro |
What is the effect of scaling the number of intermediate tasks during training on the zero-shot cross-lingual transfer performance of multil…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-21 13:07 |
| e7768b53-eb21-4a… | Gate 3 | formula_repro |
Does multimodal fine-tuning (e.g., using images alongside text) on intermediate tasks improve zero-shot cross-lingual performance on XTREME …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.5/10 | 2026-07-21 13:07 |
| b2efa482-5894-41… | Gate 3 | formula_repro |
How does intermediate-task training compare to multilingual pretraining (e.g., mT5, XLM-R) in improving zero-shot cross-lingual accuracy on …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-07-21 13:07 |
| ad22185d-f093-4d… | Gate 3 | formula_repro |
Does the choice of non-English intermediate task (e.g., NLI vs. QA) affect the zero-shot cross-lingual transfer performance more significant…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 6.5/10
|
6.1/10 | 2026-07-21 13:07 |
| 3408ff2b-e37a-41… | Gate 3 | formula_repro |
What is the impact of model size (e.g., 1B vs. 7B parameters) on zero-shot cross-lingual transfer performance when fine-tuned on non-English…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-07-21 13:07 |
| 7b818b49-216a-4f… | Gate 3 | formula_repro |
Can multimodal intermediate tasks (e.g., image-to-text or vision-language reasoning) enhance zero-shot cross-lingual performance on XTREME-R…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.4/10 | 2026-07-21 13:07 |
| a05935e0-13a8-40… | Gate 3 | formula_repro |
What is the impact of scaling the size of the pretrained model on zero-shot cross-lingual transfer performance when using English intermedia…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.0/10 · REPLICATION ATTACKER: 0.0/10
|
4.8/10 | 2026-07-21 07:02 |
| 02d7d962-188b-40… | Gate 3 | formula_repro |
How does the choice of intermediate-task difficulty (e.g., easy vs. hard tasks) impact the zero-shot cross-lingual transfer performance on X…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.2/10 | 2026-07-21 07:02 |
| 2434406b-8dd9-4f… | Gate 3 | formula_repro |
What is the impact of scaling model size on the trade-off between accuracy and inference speed when applying English intermediate-task train…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.8/10 | 2026-07-21 07:02 |
| 91d23258-d3ab-4c… | Gate 3 | formula_repro |
What is the effect of multimodal fine-tuning (audio-text) on the alignment capabilities of self-supervised speech models pre-trained on Flem…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-07-21 07:02 |
| d64a689b-9c1d-48… | Gate 3 | formula_repro |
How does the choice of intermediate-task difficulty or dataset size impact the zero-shot cross-lingual transfer performance on XTREME-R when…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 6.5/10
|
4.6/10 | 2026-07-21 07:02 |
| 98a5e5b2-d271-4b… | Gate 3 | formula_repro |
How does the performance of multilingual intermediate-task training on XTREME classification compare to English-only intermediate-task train…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-21 07:00 |
| 001d1adc-761e-45… | Gate 3 | formula_repro |
What is the impact of scaling the number of intermediate tasks (from 5 to 15) on the zero-shot cross-lingual transfer performance (measured …
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-21 00:59 |
| a4628bda-5dbd-4f… | Gate 3 | formula_repro |
How does varying the number of intermediate English tasks (1 vs. 3 vs. 9) in multi-task fine-tuning impact zero-shot performance on XTREME-R…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.5/10 | 2026-07-21 00:59 |
| c570d75c-cbcb-40… | Gate 3 | formula_repro |
How does the performance of intermediate-task training compare to direct fine-tuning on the target task when evaluated on the XTREME-R bench…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-07-21 00:59 |
| cb22d962-a3bd-4d… | Gate 3 | formula_repro |
How does the choice of intermediate language understanding tasks affect the robustness of zero-shot cross-lingual transfer in the OFA model,…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.4/10 | 2026-07-21 00:57 |
| d45dabb2-20e7-43… | Gate 3 | formula_repro |
Does scaling the number of English intermediate tasks proportionally improve zero-shot cross-lingual transfer on XTREME-R, and what is the s…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-20 18:55 |
| 31b3ecbf-521b-41… | Gate 3 | formula_repro |
How does the alignment of intermediate tasks with target task domains (e.g., natural language inference tasks for NLI benchmarks) affect the…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 6.5/10
|
5.7/10 | 2026-07-20 18:55 |
| f0f5f6e8-3424-44… | Gate 3 | formula_repro |
What is the impact of varying the size of intermediate language understanding tasks on the zero-shot cross-lingual transfer performance of t…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.4/10 | 2026-07-20 18:54 |
| ddd56d95-38a7-47… | Gate 3 | formula_repro |
How does the performance of self-supervised speech models pre-trained on mixed English and Flemish Dutch data compare to models pre-trained …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.7/10 | 2026-07-20 18:48 |
| 529194cb-654d-47… | Gate 2 | unknown | How does the trade-off between model size (100M vs 1B parameters) and intermediate-task diversity affect zero-shot accuracy on XTREME-R task… | - | 2026-07-20 15:36 |
| c985abe7-33f3-41… | Gate 3 | formula_repro |
How does the performance of multilingual language models on XTREME benchmark tasks compare when trained with intermediate English tasks vers…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.1/10 | 2026-07-20 12:12 |
| a09ecafd-49d5-4f… | Gate 3 | formula_repro |
What is the impact of scaling model size on the effectiveness of English-only intermediate-task training for zero-shot cross-lingual transfe…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-20 12:08 |
| 48432bb4-de09-4a… | Gate 3 | formula_repro |
How does scaling the number of intermediate language-understanding tasks affect zero-shot cross-lingual performance on the XTREME-R benchmar…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
8.6/10 | 2026-07-20 12:07 |
| 44daf5ec-4640-41… | Gate 3 | formula_repro |
Does pretraining on code-focused tasks (e.g., code summarization, defect detection) with models like CodeBERT improve zero-shot cross-lingua…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.6/10 | 2026-07-20 12:04 |
| 31e3f646-361e-4c… | Gate 3 | formula_repro |
What is the effect of task similarity between intermediate and target tasks on zero-shot cross-lingual performance in XTREME-R, measured by …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.5/10 | 2026-07-20 12:02 |
| 12af7721-f23f-49… | Gate 3 | formula_repro |
What is the impact of combining intermediate language-understanding tasks and code-specific tasks sequentially on the zero-shot cross-lingua…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
4.7/10 | 2026-07-20 12:00 |
| 1e064f95-767c-43… | Gate 3 | formula_repro |
How does the choice of intermediate task difficulty (e.g., easy vs. complex) affect zero-shot cross-lingual accuracy on XTREME-R when using …
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
6.2/10 | 2026-07-20 12:00 |
| 1f101ed3-92dd-4f… | Gate 3 | formula_repro |
What is the impact of scaling model size (e.g., from 7B to 30B parameters) on the performance gains from English intermediate-task training …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-20 11:56 |
| 552b42b8-fe70-4f… | Gate 3 | formula_repro |
Does scaling the size of intermediate-task datasets for English fine-tuning improve zero-shot cross-lingual performance on the XTREME-R benc…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-20 11:51 |
| bc334cb3-7016-49… | Gate 3 | formula_repro |
What is the impact of scaling up the size of the target task dataset in non-English languages on the transferability of English intermediate…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-07-20 05:48 |
| 131e8470-8d15-48… | Gate 3 | formula_repro |
How does intermediate-task training on English reasoning datasets affect the zero-shot performance of XLM-R on non-English logical reasoning…
COUNTEREXAMPLE HUNTER: 7.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.0/10 | 2026-07-20 05:48 |
| 9bcfa040-ce2e-41… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on XLM-R diminish when the target non-English tasks in XTREME-R involve co…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-07-20 05:48 |
| b59dbc8e-8ca0-4d… | Gate 3 | formula_repro |
To what extent does English intermediate-task training degrade or improve logical reasoning accuracy in code generation tasks within the XTR…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 8.7/10 · REPLICATION ATTACKER: 7.2/10
|
5.3/10 | 2026-07-20 05:48 |
| 3f997b0d-5f95-44… | Gate 3 | formula_repro |
What is the effect of using intermediate tasks from multiple languages (e.g., English and Spanish) instead of just English on the downstream…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
4.7/10 | 2026-07-20 05:48 |
| 95515558-2eaa-41… | Gate 3 | formula_repro |
What is the impact of scaling the number of languages in pre-training on the zero-shot cross-lingual retrieval performance of hybrid-trained…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 3.5/10
|
6.5/10 | 2026-07-20 05:48 |
| f33516d8-6505-4e… | Gate 3 | formula_repro |
How does scaling the amount of Flemish Dutch pre-training data impact the Word Error Rate (WER) performance of self-supervised models on Lib…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.5/10 | 2026-07-20 05:47 |
| 54200533-19a9-43… | Gate 3 | formula_repro |
What is the impact of varying the size of the bilingual lexicon used to generate code-switched training data on the inference efficiency of …
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-07-19 23:47 |
| 3ee81c97-f290-46… | Gate 3 | formula_repro |
How do cross-lingual retrieval models trained on artificially code-switched data perform on the XNLI benchmark for cross-lingual natural lan…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
8.7/10 | 2026-07-19 23:47 |
| f8e62d23-0151-47… | Gate 3 | formula_repro |
What is the computational efficiency trade-off between Flamingo and projection-based methods when scaling to 10+ low-resource languages in c…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 4.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.8/10 | 2026-07-19 23:45 |
| 784180d6-3ac2-45… | Gate 3 | formula_repro |
Does intermediate-task training on English code generation datasets improve zero-shot performance on non-English programming tasks in the XT…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-07-19 23:45 |
| 15008098-5907-46… | Gate 3 | formula_repro |
How does the intermediate-task training performance gain scale with model size when evaluated on XTREME-R, and what is the optimal model siz…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.3/10 | 2026-07-19 23:45 |
| 82a31c5b-78ae-4a… | Gate 3 | formula_repro |
How does the computational efficiency (training time, memory usage) of multi-task intermediate training scale with the number of tasks, and …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-07-19 23:45 |
| e9d9755f-6a4b-44… | Gate 3 | formula_repro |
How does intermediate-task training on non-English high-resource languages compare to multilingual intermediate-task training (e.g., XLM-R o…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-19 17:45 |
| bf458030-9bd9-46… | Gate 3 | formula_repro |
How does the combination of multiple intermediate tasks (e.g., NLI, QA, NER) affect zero-shot cross-lingual transfer in multilingual models …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-19 17:45 |
| 9e5a3ad9-e7bc-48… | Gate 3 | formula_repro |
To what extent does scaling the size of the pretrained model affect the transferability gains from multimodal intermediate-task training ver…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.8/10 | 2026-07-19 17:44 |
| fc72fbbd-4bb2-41… | Gate 3 | formula_repro |
Does intermediate-task training on structured reasoning datasets improve zero-shot transfer to non-English reasoning tasks in the XCOPA subs…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-07-19 17:44 |
| 81333cfb-67ed-43… | Gate 3 | formula_repro |
How does intermediate-task training on diverse English NLI and QA datasets affect zero-shot performance on the XTREME-R benchmark compared t…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-07-19 17:42 |
| 91c9566f-5e9a-46… | Gate 3 | formula_repro |
How does intermediate-task training on English code generation tasks influence the zero-shot cross-lingual performance on HumanEval-X for co…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-07-19 11:40 |
| 4ce2262e-808a-4b… | Gate 3 | formula_repro |
How does the choice of intermediate English task (e.g., NLI vs. QA) affect zero-shot cross-lingual transfer performance on XTREME-R as measu…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-07-19 11:39 |
| fa0472fe-8151-4b… | Gate 3 | formula_repro |
Can multimodal intermediate tasks (e.g., image-text alignment) improve zero-shot cross-lingual transfer performance on XTREME-R compared to …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 3.5/10
|
6.1/10 | 2026-07-19 11:39 |
| 40880adf-e040-42… | Gate 3 | formula_repro |
Does English intermediate-task training improve inference efficiency in multilingual models when evaluated on the XTREME-R benchmark, and ho…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-07-19 11:39 |
| 0c99e200-5b97-4c… | Gate 3 | formula_repro |
How does the order of intermediate language-understanding tasks affect zero-shot cross-lingual performance in XTREME-R, and can an optimal s…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-07-19 11:39 |
| dd1d6c9d-71a0-45… | Gate 3 | formula_repro |
Does the diversity of English intermediate tasks correlate with improved generalization scores on zero-shot cross-lingual transfer tasks wit…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.4/10 | 2026-07-19 11:39 |
| 4f908557-b44b-48… | Gate 3 | formula_repro |
How does the performance of models trained with different intermediate task sizes compare on the XTREME benchmark when evaluated for zero-sh…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.2/10
|
6.0/10 | 2026-07-19 11:39 |
| c56c5c90-fc84-4f… | Gate 3 | formula_repro |
Does training on artificially code-switched data improve the robustness of zero-shot retrieval models against domain shifts in non-English l…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.8/10 | 2026-07-19 05:37 |
| 1c2b8653-defa-49… | Gate 3 | formula_repro |
What is the effect of scaling model parameters on the marginal gains provided by English intermediate-task training for zero-shot cross-ling…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.0/10
|
6.2/10 | 2026-07-19 05:37 |
| f1309b36-0f18-4c… | Gate 3 | formula_repro |
How does the impact of English intermediate-task training on zero-shot cross-lingual reasoning in mT5-Large compare when using XTREME-U inte…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.3/10 · REPLICATION ATTACKER: 7.2/10
|
5.7/10 | 2026-07-19 05:37 |
| dab2a0fc-6f56-47… | Gate 3 | formula_repro |
How does the F1 score on XTREME-R tasks vary when using different English intermediate tasks (e.g., NLI, QA, or classification) for transfer…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.0/10 | 2026-07-19 05:37 |
| 58f87a19-411f-47… | Gate 3 | formula_repro |
What is the impact of varying the scale of intermediate-task datasets on zero-shot cross-lingual performance in XTREME-R, and how does this …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.2/10 | 2026-07-19 05:35 |
| e1c9adf2-3f96-45… | Gate 3 | formula_repro |
How does the inclusion of code-switched data in training affect the robustness of zero-shot cross-lingual retrieval models to out-of-domain …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-18 23:35 |
| ebebbf16-0a0c-45… | Gate 3 | formula_repro |
What is the impact of increasing the scale of artificially code-switched training data on the zero-shot cross-lingual retrieval performance …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 3.5/10
|
6.5/10 | 2026-07-18 23:34 |
| 9c7e1d3f-5c52-43… | Gate 3 | formula_repro |
How does the MRR@10 performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare to models uti…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.5/10 | 2026-07-18 23:34 |
| 15e54392-c51b-43… | Gate 3 | formula_repro |
What is the impact of varying the ratio of code-switched tokens in training data on the MRR@10 performance of zero-shot cross-lingual retrie…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-18 23:34 |
| 1514421e-cdf1-41… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare to models trained on…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.5/10 | 2026-07-18 23:34 |
| b1baa1a3-209c-46… | Gate 3 | formula_repro |
Does scaling the number of English intermediate tasks from one to nine yield diminishing returns in zero-shot cross-lingual transfer accurac…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-18 23:32 |
| 5038347a-ab4b-4e… | Gate 3 | formula_repro |
Does the scaling of model parameters alter the effectiveness of English intermediate-task training on zero-shot cross-lingual performance wi…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.2/10
|
5.6/10 | 2026-07-18 23:32 |
| 416c07ee-d758-4c… | Gate 3 | formula_repro |
What are the trade-offs between execution accuracy and inference throughput when using multimodal intermediate-task training versus traditio…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 6.5/10
|
7.1/10 | 2026-07-18 17:30 |
| a8143462-5810-4e… | Gate 3 | formula_repro |
What is the impact of scaling the size of the pretrained model (e.g., from 7B to 70B parameters) on the zero-shot cross-lingual transfer per…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-18 17:30 |
| 4c4173b8-6be2-45… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on XTREME-R correlate with the linguistic distance between English and the…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 6.5/10
|
6.1/10 | 2026-07-18 17:30 |
| b5e121f5-4487-47… | Gate 3 | formula_repro |
What is the impact of using intermediate tasks from different domains (e.g., natural language inference, question answering) on the zero-sho…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.0/10 | 2026-07-18 17:30 |
| 03d5c67d-8656-4f… | Gate 3 | formula_repro |
Does the scaling of model size from 1B to 70B parameters alter the efficacy of English intermediate-task training for zero-shot transfer on …
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
4.9/10 | 2026-07-18 11:29 |
| b4e23134-b09f-4f… | Gate 3 | formula_repro |
Does intermediate-task training on instruction-tuned models improve zero-shot performance on the XTREME-R benchmark compared to standard fin…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-07-18 11:28 |
| 10071122-1da5-48… | Gate 3 | formula_repro |
Does intermediate-task training on English code generation tasks improve zero-shot cross-lingual performance on structured prediction tasks …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.2/10
|
5.7/10 | 2026-07-18 11:28 |
| d722a5fa-1654-48… | Gate 3 | formula_repro |
How does English intermediate-task training impact the alignment and robustness of multilingual models in zero-shot cross-lingual settings w…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
4.7/10 | 2026-07-18 11:28 |
| 6c9540ca-8e65-42… | Gate 3 | formula_repro |
Does increasing parameter scale from 7B to 70B improve adversarial robustness scores on the XTREME-R benchmark after English intermediate-ta…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.3/10 | 2026-07-18 11:21 |
| 02199614-2b65-44… | Gate 3 | formula_repro |
What is the impact of intermediate-task training on inference efficiency (e.g., latency, throughput) when deploying cross-lingual models on …
COUNTEREXAMPLE HUNTER: 9.2/10 · CITATION AUDITOR: 5.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.4/10 | 2026-07-18 05:21 |
| 53365198-df9b-44… | Gate 3 | formula_repro |
Can multimodal intermediate-task training (e.g., vision-language tasks) further improve zero-shot cross-lingual transfer on XTREME over Engl…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-07-18 05:21 |
| 809a8818-d536-40… | Gate 3 | formula_repro |
What is the impact of scaling model size on the efficacy of English intermediate-task training for cross-lingual transfer across the XTREME …
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-07-18 05:21 |
| 2874b4c2-2beb-4d… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training persist when evaluated on downstream tasks in low-resource languages using…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-18 05:21 |
| 44a4ac89-fa9f-43… | Gate 3 | formula_repro |
How does the scaling of intermediate-task dataset size impact the accuracy-efficiency trade-off for few-shot cross-lingual transfer in XTREM…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-18 05:19 |
| 47dce42b-32bb-4b… | Gate 3 | formula_repro |
How does scaling the number of English intermediate tasks affect zero-shot performance on the XTREME-R benchmark compared to single-task int…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.3/10 · REPLICATION ATTACKER: 6.5/10
|
5.4/10 | 2026-07-18 05:19 |
| 784dc383-e33d-4b… | Gate 3 | formula_repro |
How does the incorporation of multilingual intermediate tasks (e.g., XNLI, PAWS-X) into the training pipeline affect zero-shot cross-lingual…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-18 05:17 |
| a652f228-92dc-44… | Gate 3 | formula_repro |
What is the correlation between intermediate-task training duration and zero-shot accuracy gains on non-English subsets of the XTREME benchm…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.2/10
|
5.6/10 | 2026-07-17 23:17 |
| 0857d1b1-17ed-4f… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on XTREME-R persist when the base model is scaled from 1B to 10B+ paramete…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-17 23:17 |
| 7ab66027-aefb-4b… | Gate 3 | formula_repro |
How does the choice of intermediate task (NLI vs. QA vs. NER) affect the zero-shot accuracy of mT5 and XGLM on the XTREME-R benchmark compar…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 4.2/10
|
6.4/10 | 2026-07-17 23:14 |
| f45d81f5-37dc-46… | Gate 3 | formula_repro |
Does fine-tuning a pretrained model on a mix of English and high-resource non-English intermediate tasks (e.g., 5 English + 5 non-English) i…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.7/10 | 2026-07-17 23:14 |
| 85874ead-b8c6-4e… | Gate 3 | formula_repro |
Does scaling the number of high-resource non-English intermediate tasks (e.g., 5 vs. 10 tasks) lead to improved zero-shot cross-lingual perf…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 6.3/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-07-17 23:12 |
| b34a6a89-729c-4d… | Gate 3 | formula_repro |
How does the choice of intermediate task selection strategy (random vs. diversity-optimized) affect the zero-shot cross-lingual transfer acc…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-17 17:12 |
| 0245ddab-fc9e-45… | Gate 3 | formula_repro |
How does scaling the size of intermediate language-understanding tasks (e.g., increasing the number of tasks or their complexity) affect the…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.0/10 · REPLICATION ATTACKER: 6.5/10
|
6.3/10 | 2026-07-17 17:12 |
| 606a20bc-520f-49… | Gate 3 | formula_repro |
Does intermediate-task training on reasoning-focused datasets improve zero-shot transfer performance on logical inference subsets of XTREME-…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.5/10 | 2026-07-17 17:12 |
| 455b2ebc-e5b9-47… | Gate 3 | formula_repro |
What is the impact of scaling the size of English intermediate-task datasets on the zero-shot cross-lingual transfer accuracy for multilingu…
COUNTEREXAMPLE HUNTER: 3.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-17 17:12 |
| 390f8c8a-9981-4f… | Gate 3 | formula_repro |
To what extent does increasing model size mitigate the performance drop of English intermediate-task training when transferring to syntactic…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-07-17 17:01 |
| 4b05b471-61ba-42… | Gate 3 | formula_repro |
How does scaling model parameters from 3B to 13B affect the marginal gain of English intermediate-task training on XTREME classification acc…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.3/10 | 2026-07-17 16:59 |
| 3a7bed66-f97e-4e… | Gate 3 | formula_repro |
Does English intermediate-task training on reasoning-heavy datasets improve zero-shot performance on logical inference tasks within XTREME-R…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.3/10 | 2026-07-17 10:57 |
| 88dfdda5-dd5d-41… | Gate 3 | formula_repro |
What is the correlation between the size of the English intermediate task dataset and the zero-shot transfer accuracy on XTREME question-ans…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
8.6/10 | 2026-07-17 10:57 |
| 63d2ba58-a4e0-47… | Gate 3 | formula_repro |
What is the effect of scaling the number of intermediate language understanding tasks on the inference efficiency and zero-shot performance …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.1/10 | 2026-07-17 10:55 |
| 1628a25b-248c-46… | Gate 3 | formula_repro |
How does intermediate-task training on high-resource non-English datasets impact zero-shot accuracy on XTREME-R compared to English intermed…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-17 10:51 |
| cd670ca4-c227-49… | Gate 3 | formula_repro |
Does intermediate-task training on English reasoning datasets improve zero-shot transfer accuracy on logical inference subsets of the XTREME…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-17 10:47 |
| 9b0d937f-1e7a-45… | Gate 3 | formula_repro |
How does the alignment of intermediate tasks with target-language domains (e.g., legal, medical) impact the accuracy of zero-shot cross-ling…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-17 10:43 |
| 68074e40-ffc9-4f… | Gate 3 | formula_repro |
How does the intermediate-task fine-tuning order (sequential vs. concurrent) affect zero-shot cross-lingual transfer accuracy on XTREME-R be…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-17 10:42 |
| c69933d0-ace9-4e… | Gate 3 | formula_repro |
How does the effectiveness of intermediate-task training on English for structured prediction tasks in XTREME-R vary when using few-shot lea…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
5.6/10 | 2026-07-17 04:40 |
| 96588404-50bf-43… | Gate 3 | formula_repro |
Does intermediate-task training on English code generation datasets improve zero-shot cross-lingual transfer performance on multilingual pro…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-07-17 04:38 |
| c936bc0c-2c32-43… | Gate 3 | formula_repro |
How does the performance of multilingual intermediate-task training compare to monolingual English intermediate-task training in zero-shot c…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 9.2/10
|
6.7/10 | 2026-07-17 04:38 |
| 182478f9-8ca3-4a… | Gate 3 | formula_repro |
Does scaling the intermediate-task dataset size proportionally to the model size improve zero-shot F1 scores on XTREME classification tasks …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.3/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-17 04:36 |
| 69b372a8-ef30-40… | Gate 3 | formula_repro |
What is the efficiency trade-off between intermediate-task training and direct fine-tuning for zero-shot cross-lingual transfer on XTREME-R,…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.0/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-07-17 04:29 |
| 16af23c6-3028-46… | Gate 3 | formula_repro |
Does the performance gain from intermediate-task training in zero-shot cross-lingual transfer persist when scaling to larger, state-of-the-a…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-07-17 04:29 |
| ff379c72-0e23-4c… | Gate 3 | formula_repro |
How does intermediate-task training on code generation datasets like HumanEval impact zero-shot performance on the programming subset of XTR…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-07-17 04:27 |
| 64378984-239e-4d… | Gate 3 | formula_repro |
How does domain-specific intermediate-task training (e.g., code, reasoning) influence the zero-shot cross-lingual transfer performance of mu…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.7/10 | 2026-07-16 22:25 |
| cac34759-415b-49… | Gate 3 | formula_repro |
Does pre-training dense retrievers on code-switched data improve zero-shot cross-lingual transfer performance on multilingual benchmarks lik…
COUNTEREXAMPLE HUNTER: 8.2/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 7.5/10
|
8.1/10 | 2026-07-16 22:23 |
| 9bd94372-89c8-42… | Gate 3 | formula_repro |
Does intermediate-task training on multilingual tasks (e.g., XNLI, TyDi QA) improve zero-shot cross-lingual transfer performance on low-reso…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-16 22:23 |
| 0c005fb7-459e-42… | Gate 3 | formula_repro |
What is the effect of mixed intermediate-task training (multiple tasks sequentially) on the inference efficiency of XLM-R Large for zero-sho…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.0/10 | 2026-07-16 16:23 |
| 5ec4d212-b29c-46… | Gate 3 | formula_repro |
To what extent does the choice of English intermediate task (e.g., NLI, QA) influence the zero-shot cross-domain generalization of multiling…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.0/10 · REPLICATION ATTACKER: 7.5/10
|
6.3/10 | 2026-07-16 16:23 |
| f36bea13-557c-4a… | Gate 3 | formula_repro |
What is the impact of using multimodal intermediate tasks (e.g., image-captioning) before fine-tuning on zero-shot cross-lingual accuracy in…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.1/10 | 2026-07-16 16:22 |
| 9589d524-f845-45… | Gate 3 | formula_repro |
Does intermediate-task training on English reasoning datasets improve zero-shot cross-lingual logical inference accuracy on the XCOPA benchm…
COUNTEREXAMPLE HUNTER: 7.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-16 16:22 |
| 444830c2-63fa-47… | Gate 3 | formula_repro |
Can intermediate-task training with English tasks improve the reasoning capabilities of multilingual LLMs on XTREME-R, as measured by the MM…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-16 16:11 |
| 10ff1033-802b-46… | Gate 3 | formula_repro |
How does the performance of intermediate-task training with multiple English tasks compare to single-task fine-tuning on the XTREME-R benchm…
COUNTEREXAMPLE HUNTER: 8.0/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.0/10
|
6.2/10 | 2026-07-16 16:11 |
| 40350efc-708d-44… | Gate 3 | formula_repro |
Does the order or sequence of intermediate English tasks in multi-task fine-tuning affect the zero-shot cross-lingual transfer performance o…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-07-16 16:11 |
| 65bd52bc-d0df-49… | Gate 3 | formula_repro |
Does incorporating multimodal intermediate tasks prior to target fine-tuning enhance zero-shot cross-lingual transfer scores on the XTREME b…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-07-16 10:09 |
| e88fc853-c3ab-47… | Gate 3 | formula_repro |
How does intermediate-task training on code generation datasets affect zero-shot transfer performance on natural language understanding task…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.4/10 | 2026-07-16 10:08 |
| f78e6bb4-47c0-49… | Gate 3 | formula_repro |
Does scaling the number of non-English intermediate tasks during fine-tuning improve zero-shot accuracy on the XTREME benchmark compared to …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.4/10 | 2026-07-16 10:08 |
| 2176a52f-cd9c-41… | Gate 3 | formula_repro |
Does intermediate-task training on English legal benchmarks like LIMLab improve zero-shot cross-lingual F1 scores on XTREME-R compared to ge…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.2/10 | 2026-07-16 10:04 |
| 4c3a4051-849f-49… | Gate 3 | formula_repro |
How does the choice of intermediate task difficulty (e.g., complexity of language understanding tasks) impact the zero-shot cross-lingual tr…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.3/10
|
4.9/10 | 2026-07-16 04:04 |
| 32dabe11-5b29-4c… | Gate 3 | formula_repro |
Does intermediate-task training on English datasets improve reasoning capabilities on non-English subsets of the XTREME benchmark compared t…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-16 04:04 |
| d32c5302-0d7c-4c… | Gate 3 | formula_repro |
Does increasing the number of intermediate tasks in multitask fine-tuning improve zero-shot cross-lingual transfer accuracy on XTREME-R beyo…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-16 04:04 |
| 53f29c20-3a48-46… | Gate 3 | formula_repro |
To what extent does English intermediate-task training impact the convergence speed and GPU memory footprint of multilingual transformers on…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.8/10 | 2026-07-16 04:04 |
| 7ca6d7ce-9d84-48… | Gate 3 | formula_repro |
Does increasing the parameter scale of the backbone model diminish the relative performance gain provided by English intermediate-task train…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-16 04:04 |
| b0edc320-3ba9-4e… | Gate 3 | formula_repro |
Does English intermediate-task training on reasoning-heavy datasets improve logical inference scores on non-English subsets of XTREME more e…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-07-16 04:02 |
| 7da50001-ade4-41… | Gate 3 | formula_repro |
Does intermediate-task training on a diverse set of English tasks improve zero-shot cross-lingual performance on XTREME more than training o…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-15 22:02 |
| 77dd4037-fb53-49… | Gate 3 | formula_repro |
How does intermediate-task training on CodeXGLUE benchmarks impact zero-shot performance on non-English natural language inference tasks in …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-15 22:01 |
| b5080d0e-5b5b-44… | Gate 3 | formula_repro |
How does the choice of English intermediate-task difficulty (e.g., complexity of the task, dataset size) affect zero-shot cross-lingual tran…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-07-15 22:01 |
| b0ff9dcc-6f89-49… | Gate 3 | formula_repro |
How does scaling the intermediate-task dataset size affect the zero-shot cross-lingual transfer performance on XTREME-R logical inference ta…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 6.5/10
|
7.6/10 | 2026-07-15 21:59 |
| 42e98ae0-bb91-46… | Gate 3 | formula_repro |
How does the hybrid batch training strategy compare to continuous batch training in terms of zero-shot retrieval accuracy on the XQuAD bench…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-07-15 15:55 |
| bec7acfa-4bf7-44… | Gate 3 | formula_repro |
How does the inference efficiency of dense retrievers, when trained on artificially code-switched data versus standard monolingual data, com…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 6.5/10
|
4.3/10 | 2026-07-15 15:54 |
| 45258043-0243-48… | Gate 3 | formula_repro |
Does the simultaneous optimization of monolingual and cross-lingual objectives in hybrid batch training degrade performance on domain-specif…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-15 15:54 |
| b16c2f48-c168-4e… | Gate 3 | formula_repro |
Does intermediate-task training on multimodal models (e.g., CLIP) improve zero-shot XTREME-R performance compared to text-only models when e…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-07-15 09:50 |
| 1ff6ed3e-dd7e-4c… | Gate 3 | formula_repro |
How does increasing the size of English intermediate-task training datasets impact zero-shot accuracy on the XTREME benchmark for typologica…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-15 09:50 |
| 81fc3347-a788-49… | Gate 3 | formula_repro |
How does intermediate-task training on high-resource English datasets affect zero-shot performance on XTREME-R reasoning tasks compared to t…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
4.7/10 | 2026-07-15 09:49 |
| ba85339b-8c74-4d… | Gate 3 | formula_repro |
Do Flemish Dutch self-supervised speech models (e.g., HuBERT or Conformer) trained on varying amounts of unlabeled data show consistent impr…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 3.0/10
|
5.7/10 | 2026-07-15 03:45 |
| c27c706b-7c12-41… | Gate 3 | formula_repro |
To what extent does simultaneous optimization for monolingual and cross-lingual objectives affect zero-shot performance on code-related retr…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.3/10 | 2026-07-15 03:45 |
| 77592950-5892-44… | Gate 3 | formula_repro |
What is the impact of scaling the size of multimodal datasets (e.g., Flickr30K, COCO) on the zero-shot retrieval performance of multilingual…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-07-15 03:44 |
| 46196291-d0be-4a… | Gate 2 | unknown | What is the effect of the hybrid batch training strategy on the inference throughput (queries/second) of zero-shot cross-lingual retrieval m… | - | 2026-07-15 02:55 |
| d4b97c17-4a9e-49… | Gate 3 | formula_repro |
How does varying the ratio of cross-lingual to monolingual training batches impact zero-shot retrieval accuracy on the XTYLE benchmark for l…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.0/10
|
5.2/10 | 2026-07-14 21:28 |
| cf5d187c-02e5-46… | Gate 3 | formula_repro |
How does the hybrid batch training strategy impact zero-shot cross-lingual retrieval performance when scaled to larger multilingual language…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 2.5/10
|
6.9/10 | 2026-07-14 15:02 |
| fdfe9203-da53-46… | Gate 3 | formula_repro |
What is the impact of varying the ratio of monolingual, cross-lingual, and multilingual data in the hybrid batch training approach on the re…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-07-14 15:00 |
| 74c1f45d-8a8e-40… | Gate 3 | formula_repro |
How does scaling the size of the hybrid batch affect the trade-off between monolingual and cross-lingual retrieval performance on MIRACL, an…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 2.0/10
|
6.8/10 | 2026-07-14 14:58 |
| 1662dee5-ecbb-46… | Gate 3 | formula_repro |
How does the scaling of the hybrid batch training strategy with increasing dataset size affect zero-shot cross-lingual retrieval performance…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 0.0/10
|
6.2/10 | 2026-07-14 14:49 |
| 74bb1055-6921-4e… | Gate 3 | formula_repro |
How does the synergistic hybrid batch training strategy compare to existing monolingual and multilingual pre-training methods in terms of ze…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 2.0/10
|
6.5/10 | 2026-07-14 14:46 |
| 72bc8aa3-a972-40… | Gate 3 | formula_repro |
How does hybrid batch training affect zero-shot cross-lingual retrieval accuracy on the BEIR benchmark for 7B+ parameter models compared to …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.0/10
|
9.0/10 | 2026-07-14 08:18 |
| a0df1ece-192c-44… | Gate 3 | formula_repro |
How robust are multimodal projection-based methods for cross-lingual NER in low-resource languages when evaluated across different domains (…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.1/10 | 2026-07-14 08:14 |
| 435fc697-7305-4b… | Gate 3 | formula_repro |
How does the performance of projection-based cross-lingual NER compare to few-shot learning with multilingual LLMs on XTREME-R low-resource …
COUNTEREXAMPLE HUNTER: 10.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 2.0/10
|
7.2/10 | 2026-07-14 08:13 |
| 4595b97f-6387-47… | Gate 3 | formula_repro |
Does the synergistic training strategy for monolingual and cross-lingual retrieval maintain robustness against language imbalance in low-res…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 9.0/10
|
6.7/10 | 2026-07-14 08:12 |
| 17a4b8f1-74c9-47… | Gate 3 | formula_repro |
How does the choice of English intermediate-task complexity (e.g., XNLI vs. PAWS) affect zero-shot cross-lingual transfer performance on XTR…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 6.2/10
|
8.1/10 | 2026-07-14 02:02 |
| 217b52aa-3a99-4b… | Gate 3 | formula_repro |
How does intermediate-task training on multiple non-English source languages affect zero-shot transfer performance compared to English-only …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.3/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-14 02:00 |
| 752b5eef-7abe-4e… | Gate 3 | formula_repro |
What is the impact of domain-specific intermediate-task training (e.g., legal or scientific tasks) on zero-shot cross-lingual transfer perfo…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
5.6/10 | 2026-07-14 02:00 |
| ca35de24-f8da-4c… | Gate 3 | formula_repro |
What is the effect of varying the size of the pretrained model (e.g., base vs. large) on zero-shot cross-lingual transfer accuracy when usin…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.8/10 | 2026-07-14 01:58 |
| ce3eff31-f6a6-43… | Gate 3 | formula_repro |
Does the diversity of intermediate English tasks correlate with improved robustness in zero-shot cross-lingual transfer for syntactic analys…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
8.6/10 | 2026-07-14 01:58 |
| 197b3f97-3393-41… | Gate 3 | formula_repro |
To what extent does English intermediate-task training degrade zero-shot performance on low-resource languages within the XTREME benchmark c…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
4.9/10 | 2026-07-14 01:54 |
| 17320b22-9b66-4c… | Gate 3 | formula_repro |
How does intermediate-task training on non-English source languages compare to English-only intermediate training for zero-shot transfer per…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.5/10
|
6.2/10 | 2026-07-14 01:54 |
| 68ffd512-70d4-48… | Gate 3 | formula_repro |
What is the impact of multilingual intermediate-task training (e.g., using XTREME-M) on downstream zero-shot cross-lingual transfer performa…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-13 19:51 |
| daac403a-db0b-49… | Gate 3 | formula_repro |
How does the semantic diversity of English intermediate tasks affect zero-shot cross-lingual accuracy on the XTREME-R benchmark compared to …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-13 19:50 |
| 0d26129f-0ab4-4e… | Gate 3 | formula_repro |
How does scaling pretrained model parameters from base to huge affect zero-shot accuracy on XTREME classification tasks compared to sequence…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-13 19:49 |
| 1de50785-c0f5-4f… | Gate 3 | formula_repro |
Does intermediate-task training on English question answering datasets improve zero-shot performance on non-English reasoning subsets of XTR…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.2/10 | 2026-07-13 19:46 |
| f38564a2-c4f9-40… | Gate 3 | formula_repro |
What is the comparative performance gain of English intermediate-task training versus continued pretraining on low-resource language corpora…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-13 19:44 |
| 46910b2b-025c-45… | Gate 3 | formula_repro |
Does English intermediate-task training improve the reasoning capabilities of small language models on multilingual mathematical reasoning s…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-13 19:42 |
| 9a708a07-ea33-48… | Gate 3 | formula_repro |
What is the effect of model size on the effectiveness of English intermediate-task training for zero-shot cross-lingual transfer to Finnish,…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-07-13 19:41 |
| 49c0c783-06c0-47… | Gate 3 | formula_repro |
How does the choice of intermediate task domain (e.g., NLI vs. QA) impact zero-shot cross-lingual transfer performance on Finnish, as measur…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-07-13 19:38 |
| d941fada-172f-44… | Gate 3 | formula_repro |
How does the performance of English intermediate-task training compare to multilingual intermediate-task training for zero-shot cross-lingua…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-13 19:35 |
| 2c43888a-331a-45… | Gate 3 | formula_repro |
How does English intermediate-task training impact the inference latency and throughput of mT5 models on XTREME-R classification tasks compa…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-07-13 19:31 |
| dfc53687-6eea-41… | Gate 3 | formula_repro |
What is the impact of incorporating multilingual contrastive learning objectives during hybrid batch training on the zero-shot cross-lingual…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-07-13 13:28 |
| 34701f90-56a1-4c… | Gate 3 | formula_repro |
What is the impact of using target-language-specific development sets instead of English for model selection in zero-shot cross-lingual eval…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.0/10
|
9.1/10 | 2026-07-13 13:27 |
| 2a02a4df-06c6-49… | Gate 3 | formula_repro |
What is the impact of domain adaptation (e.g., news vs. scientific domains) on the zero-shot cross-lingual retrieval performance of models t…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-07-13 13:26 |
| d316b2c1-3703-46… | Gate 3 | formula_repro |
Can multimodal pretraining (e.g., CLIP-based models) improve zero-shot cross-lingual retrieval performance when trained on code-switched dat…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.7/10 | 2026-07-13 13:25 |
| f80882c7-44f9-44… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data vary with different sizes of…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.2/10
|
5.9/10 | 2026-07-13 13:23 |
| 276a7766-fc7b-45… | Gate 3 | formula_repro |
How does the inference efficiency (measured in tokens per second) of cross-lingual retrieval models trained on code-switched data compare to…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.7/10 | 2026-07-13 13:17 |
| 43a50c16-2e44-48… | Gate 3 | formula_repro |
How does the choice of bilingual lexicon source (machine-translated vs. human-annotated) affect the performance of XLM-R on BEIR benchmark t…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-07-13 13:12 |
| 213e7c7c-d6f8-46… | Gate 2 | unknown | What is the impact of different bilingual lexicon sizes on the accuracy of artificially code-switched data for training cross-lingual retrie… | - | 2026-07-13 10:44 |
| d423bab6-3ffa-41… | Gate 3 | formula_repro |
How does the performance of XLM-R and mT5 compare on zero-shot cross-lingual logical inference tasks when intermediate-task training is cond…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 1.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-07-13 07:09 |
| 68c353a1-3709-4e… | Gate 3 | formula_repro |
How does domain-specific intermediate-task training (e.g., reasoning or code-related tasks) impact zero-shot cross-lingual transfer in XTREM…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
4.7/10 | 2026-07-13 07:09 |
| ea5bc9f9-6f77-4f… | Gate 3 | formula_repro |
How does the zero-shot cross-lingual transfer performance on XTREME-R differ when using multitask intermediate training with different combi…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-07-13 07:09 |
| 5a267bed-94ca-4f… | Gate 3 | formula_repro |
What is the trade-off between inference efficiency and cross-lingual transfer performance when applying intermediate-task training to XLM-R …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-07-13 07:09 |
| b08f7d86-7579-48… | Gate 3 | formula_repro |
Does scaling the size of the pretrained model (e.g., from 110M to 300M parameters) mitigate or amplify the benefits of intermediate-task tra…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 7.5/10
|
8.3/10 | 2026-07-13 07:08 |
| a4b85c21-e488-42… | Gate 3 | formula_repro |
Does English intermediate-task training degrade zero-shot transfer accuracy on non-Latin script languages within the XTREME benchmark compar…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.1/10 | 2026-07-13 07:07 |
| 0bb32d5c-3c65-49… | Gate 3 | formula_repro |
How does the choice of intermediate task difficulty (e.g., simple vs. complex NLI tasks) impact zero-shot cross-lingual transfer performance…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.6/10 | 2026-07-13 07:04 |
| 5fda442c-b8a7-42… | Gate 3 | formula_repro |
How does the addition of domain-specific intermediate tasks (e.g., legal or medical) affect zero-shot cross-lingual accuracy on specialized …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-07-13 07:03 |
| 748f486a-fd45-40… | Gate 3 | formula_repro |
How does intermediate-task training on non-English high-resource languages affect zero-shot transfer performance on the XTREME benchmark com…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.4/10 | 2026-07-13 01:03 |
| 8b8139d8-cde2-4f… | Gate 3 | formula_repro |
Can multimodal intermediate tasks (e.g., image captioning) enhance zero-shot cross-lingual transfer for text-only tasks in XTREME compared t…
COUNTEREXAMPLE HUNTER: 7.0/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
5.9/10 | 2026-07-13 01:03 |
| ad090a92-25d5-4d… | Gate 3 | formula_repro |
How does the choice of intermediate task complexity (e.g., NLI vs. QA vs. STS) impact zero-shot cross-lingual transfer performance on XTREME…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.5/10 | 2026-07-13 01:03 |
| 73db1de7-e4b7-45… | Gate 3 | formula_repro |
To what extent does the pre-training corpus size of the base model (e.g., mT5-Small vs. mT5-XXL) influence the effectiveness of English inte…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-07-13 01:02 |
| ac284c7c-86df-40… | Gate 3 | formula_repro |
How does the choice of English intermediate-task complexity (e.g., NLI, QA, or sentiment analysis) affect zero-shot cross-lingual transfer p…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
5.9/10 | 2026-07-13 00:59 |
| 7580a87d-0b4c-40… | Gate 3 | formula_repro |
Does intermediate-task training with multilingual models (e.g., mT5, XLM-R) improve zero-shot cross-lingual transfer on XTREME-R compared to…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-07-13 00:47 |
| 4fa507b1-d9b8-43… | Gate 3 | formula_repro |
How does the scaling of model size (e.g., 7B vs. 33B parameters) affect the effectiveness of intermediate-task training for zero-shot cross-…
COUNTEREXAMPLE HUNTER: 6.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.0/10 | 2026-07-13 00:45 |
| 1b14e099-7eda-4d… | Gate 3 | formula_repro |
What is the impact of English versus target-language intermediate-task training on the few-shot learning accuracy of XLM-R across diverse la…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-07-12 18:45 |
| 96cccc17-cd06-43… | Gate 3 | formula_repro |
How robust is the zero-shot cross-lingual transfer performance of models trained with English intermediate tasks when evaluated on adversari…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.5/10
|
7.0/10 | 2026-07-12 18:44 |
| e4286944-969c-4a… | Gate 3 | formula_repro |
Does intermediate-task training with code-related benchmarks (e.g., HumanEval or MBPP) improve zero-shot cross-lingual code reasoning in XTR…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 6.5/10
|
7.7/10 | 2026-07-12 18:43 |
| fe1b2a67-de39-4f… | Gate 3 | formula_repro |
What is the impact of intermediate-task training on English on zero-shot performance for code generation tasks in XTREME-R, and how does the…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-07-12 18:43 |
| b1b47d5c-86a8-44… | Gate 3 | formula_repro |
What is the effect of using a combination of multiple intermediate language-understanding tasks on zero-shot cross-lingual transfer accuracy…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.3/10 · REPLICATION ATTACKER: 9.2/10
|
8.2/10 | 2026-07-12 18:41 |
| 105e634f-9c10-45… | Gate 2 | unknown | Does intermediate-task training on multilingual benchmarks (e.g., XTREME-R, TyDi QA) improve zero-shot cross-lingual transfer performance of… | - | 2026-07-12 14:15 |
| 8df2d868-a591-46… | Gate 3 | formula_repro |
What is the effect of scaling the size of the intermediate-task training dataset on zero-shot cross-lingual performance in XTREME-R, and doe…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.3/10 | 2026-07-12 12:35 |
| 43efbd2b-fd9d-4f… | Gate 3 | formula_repro |
How does the impact of intermediate-task training on zero-shot cross-lingual transfer performance compare between monolingual and multilingu…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-12 12:35 |
| 5dab86d5-0353-4c… | Gate 3 | formula_repro |
How does the impact of intermediate-task training vary across different language families or typologically diverse languages in zero-shot cr…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-12 12:35 |
| 678230a2da264434… | Gate 3 | lean4_proof |
For every natural number n, the square of n modulo 4 is either 0 or 1.
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 0.0/10
|
6.0/10 | 2026-07-12 12:35 |
| 6c2fdc2b-18e1-45… | Gate 2 | unknown | How does simultaneous fine-tuning on all languages compare to sequential fine-tuning with varying numbers of intermediate languages in terms… | - | 2026-07-12 12:01 |
| 24d1d2e0-ef83-41… | Gate 3 | formula_repro |
What is the impact of using multilingual intermediate-task training (instead of English-only) on zero-shot cross-lingual transfer performanc…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-12 06:33 |
| 390d26e1-c6c2-47… | Gate 3 | formula_repro |
How does the choice of intermediate-task difficulty (e.g., easy vs. hard) affect the zero-shot cross-lingual transfer performance on XTREME …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-12 06:33 |
| 3231b724-c77d-48… | Gate 3 | formula_repro |
How does intermediate-task training on code generation datasets like HumanEval affect zero-shot cross-lingual transfer performance on the XT…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-12 06:32 |
| fd69eb95-2a53-45… | Gate 3 | formula_repro |
How does intermediate-task training on high-resource English NLI datasets affect zero-shot transfer accuracy to low-resource languages in th…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-12 06:32 |
| cba99dab-85a9-4e… | Gate 3 | formula_repro |
Does intermediate-task training on English code generation tasks improve zero-shot transfer performance on the PAWS-X dataset within XTREME-…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-12 06:32 |
| b04b32e9-ddda-48… | Gate 3 | formula_repro |
How does increasing the parameter count of multilingual pretrained models affect zero-shot XTREME-R accuracy when fine-tuned on English-only…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-07-12 00:30 |
| 1ad948e6-93cf-4e… | Gate 3 | formula_repro |
How does scaling the number of intermediate tasks influence the inference efficiency (throughput, latency) of language models on XTREME-R be…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-12 00:26 |
| 648334e1-1013-47… | Gate 3 | formula_repro |
Does the effectiveness of intermediate-task training for zero-shot cross-lingual transfer vary significantly between different language fami…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 8.5/10
|
8.1/10 | 2026-07-12 00:26 |
| 3b0e25e7-35d1-4e… | Gate 3 | formula_repro |
Can intermediate-task training on non-English monolingual tasks (e.g., Spanish, Arabic) further enhance zero-shot cross-lingual transfer to …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.8/10 | 2026-07-12 00:24 |
| 18b0ffaf-ae1e-49… | Gate 3 | formula_repro |
What is the impact of scaling the number of intermediate tasks (e.g., 3, 6, or 9 tasks) fine-tuned in the target language on zero-shot cross…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-11 18:17 |
| 762b3cf1-ffc9-40… | Gate 3 | formula_repro |
Does fine-tuning multilingual models on intermediate tasks in the target language improve zero-shot performance on XTREME-R more than fine-t…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.3/10
|
6.1/10 | 2026-07-11 18:17 |
| ac9c2cc5-2762-45… | Gate 3 | formula_repro |
Does English intermediate-task training degrade inference throughput or increase latency during zero-shot evaluation on XTREME sequence labe…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.2/10
|
7.6/10 | 2026-07-11 18:15 |
| 33f5fa65-5883-42… | Gate 3 | formula_repro |
How does cross-domain adaptation (e.g., from NLI to QA tasks) in intermediate-task training affect zero-shot cross-lingual transfer performa…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-11 18:13 |
| 027bb9c2-2162-48… | Gate 3 | formula_repro |
How does the F1 score degradation of XLM-R in sequential fine-tuning compare to simultaneous fine-tuning for idiom detection tasks across En…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 5.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.0/10 | 2026-07-11 18:11 |
| 5abed367-dfda-4d… | Gate 3 | formula_repro |
What is the impact of domain adaptation techniques on the robustness of self-supervised speech models pre-trained on Flemish Dutch when eval…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-11 18:09 |
| c3f70e70-d524-4a… | Gate 3 | formula_repro |
How does the cross-lingual transfer performance of self-supervised speech models pre-trained on Flemish Dutch compare to models pre-trained …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-11 18:08 |
| f9df9bdc-6e04-4e… | Gate 3 | formula_repro |
How does the cross-lingual transferability of self-supervised speech models trained on varying amounts of unlabeled Flemish Dutch data compa…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-07-11 18:07 |
| 602aaab7-9210-4e… | Gate 2 | unknown | How does the scaling of model size (e.g., XLM-R vs. XLM-R-Large) influence the cross-lingual transfer performance of sequential fine-tuning … | - | 2026-07-11 13:09 |
| 044768f1-f569-4f… | Gate 3 | formula_repro |
Does the simultaneous optimization strategy for monolingual and cross-lingual retrieval degrade performance on multimodal retrieval tasks wh…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 2.0/10
|
6.7/10 | 2026-07-11 12:02 |
| 4635e52a-6c6f-49… | Gate 3 | formula_repro |
Does incorporating domain-specific pre-training data alongside artificially code-switched samples improve the robustness of zero-shot cross-…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.5/10 | 2026-07-11 12:01 |
| 5e7f7307-7b2b-4b… | Gate 3 | formula_repro |
How does training on artificially code-switched data affect the zero-shot retrieval accuracy of dense retrievers on the MLDR benchmark compa…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-07-11 12:00 |
| 29c629ce-b933-43… | Gate 3 | formula_repro |
Can the hybrid batch training strategy be scaled effectively to large-scale multilingual models (e.g., LLaMA, PaLM) while maintaining zero-s…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-11 11:58 |
| 2679ae6b-a97e-48… | Gate 3 | formula_repro |
How does the incorporation of domain-specific monolingual data into hybrid batch training affect the XTREME-R score for zero-shot cross-ling…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.5/10 | 2026-07-11 11:57 |
| c086f49e-35f8-45… | Gate 3 | formula_repro |
Does the hybrid batch training strategy improve robustness to domain shifts in zero-shot cross-lingual retrieval, as measured by performance…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 2.5/10
|
6.0/10 | 2026-07-11 11:56 |
| 2a924f3a-3243-4f… | Gate 3 | formula_repro |
What is the impact of varying amounts of Flemish Dutch pre-training data on the Word Error Rate (WER) of self-supervised models when evaluat…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.0/10
|
8.5/10 | 2026-07-11 11:55 |
| 8edec612-e54f-45… | Gate 3 | formula_repro |
Can the effectiveness of zero-shot cross-lingual retrieval models trained on artificially code-switched data be improved by incorporating mu…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.0/10
|
9.1/10 | 2026-07-11 05:55 |
| 8dede7d9-8fd6-43… | Gate 3 | formula_repro |
How do multimodal language models (e.g., CLIP, Flamingo) perform in cross-lingual NER tasks compared to text-only teacher-student frameworks…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-07-11 05:54 |
| 8b3a0449-9c2d-4a… | Gate 3 | formula_repro |
How does the inclusion of multilingual data from typologically diverse languages in hybrid batch training impact the robustness of zero-shot…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 2.0/10
|
6.8/10 | 2026-07-11 05:54 |
| 69acc124-a200-47… | Gate 3 | formula_repro |
How does incorporating domain-specific monolingual data during hybrid batch training affect zero-shot retrieval performance on the XTYLE ben…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-11 05:53 |
| dee133fe-dcc1-4d… | Gate 3 | formula_repro |
How does the proposed hybrid batch training strategy affect the robustness of multilingual retrieval models on adversarial cross-lingual inp…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 2.0/10
|
6.8/10 | 2026-07-11 05:53 |
| 297aa16c-f9d3-44… | Gate 3 | formula_repro |
How does the hybrid batch training strategy perform on the XLM-R benchmark compared to specialized monolingual and cross-lingual training in…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.0/10
|
8.8/10 | 2026-07-11 05:53 |
| 02aca079-6ec2-41… | Gate 3 | formula_repro |
How does scaling the number of languages in the hybrid batch training strategy affect the generalization gap in zero-shot cross-lingual retr…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.0/10
|
9.2/10 | 2026-07-11 05:53 |
| 2f94407a-1935-4b… | Gate 3 | formula_repro |
What are the robustness differences in zero-shot cross-lingual retrieval performance between the hybrid batch training approach and monoling…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-07-11 05:52 |
| 6bd51efb-a6d7-46… | Gate 2 | unknown | What is the impact of varying the masking ratio in the masked language model on the correction accuracy of discrete speech units for accent … | - | 2026-07-11 01:29 |
| a56e9d42-f25d-46… | Gate 3 | formula_repro |
How does the hybrid batch training strategy compare to traditional multilingual pre-training methods in terms of zero-shot cross-lingual ret…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.0/10 | 2026-07-10 23:50 |
| 0faddb0d-4101-47… | Gate 3 | formula_repro |
Can the hybrid batch training strategy maintain monolingual accuracy while improving cross-lingual retrieval performance when evaluated on t…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.3/10 | 2026-07-10 23:48 |
| 53e5b162-97fc-4a… | Gate 3 | formula_repro |
How does the effectiveness of multilingual intermediate-task training compare to English-only intermediate-task training in zero-shot cross-…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.0/10
|
9.1/10 | 2026-07-10 09:42 |
| 42a552f4-034c-4a… | Gate 3 | formula_repro |
How does the choice of intermediate task domain (e.g., natural language inference vs. question answering) influence the effectiveness of Eng…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.0/10
|
9.1/10 | 2026-07-10 09:42 |
| 3b4bf387-c61a-4f… | Gate 3 | formula_repro |
How does the scale of pretrained multilingual models (e.g., 100M vs. 1B parameters) interact with cross-lingual intermediate-task training t…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-10 09:41 |
| f6c5c53d-b39a-47… | Gate 3 | formula_repro |
How does scaling the diversity of intermediate tasks in non-English high-resource languages (e.g., Spanish, French) compare to using a singl…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.1/10 · REPLICATION ATTACKER: 9.5/10
|
6.7/10 | 2026-07-10 09:41 |
| 5a649855-adde-49… | Gate 3 | formula_repro |
Does the use of multilingual intermediate tasks instead of English-only tasks improve zero-shot cross-lingual transfer performance on XTREME…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-10 09:40 |
| c2c9058d-96ba-49… | Gate 3 | formula_repro |
What is the impact of varying the number and diversity of intermediate language-understanding tasks on zero-shot cross-lingual transfer perf…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-07-10 09:39 |
| dc08282f-70b6-4d… | Gate 3 | formula_repro |
Does the performance improvement from English intermediate-task training on XTREME-R generalize to other cross-lingual benchmarks like XTREM…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-10 09:39 |
| 205c70bd-29fa-43… | Gate 3 | formula_repro |
How does the choice of intermediate task difficulty (e.g., syntactic vs. semantic tasks) affect the zero-shot cross-lingual transfer perform…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-07-10 09:38 |
| 0aab46cd-7023-48… | Gate 3 | formula_repro |
What is the impact of scaling the size of the pretrained language model on the effectiveness of intermediate-task training for zero-shot cro…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-10 09:38 |
| c2545403-a205-49… | Gate 2 | unknown | Does sequential fine-tuning on intermediate tasks improve robustness against adversarial paraphrasing in low-resource euphemism detection? | - | 2026-07-10 06:59 |
| 758ba004-9c65-43… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual transfer on the XTREME benchmark compare when using language-specific intermediate tasks…
COUNTEREXAMPLE HUNTER: 2.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.2/10
|
5.4/10 | 2026-07-10 03:36 |
| 1c610ddc-fe63-49… | Gate 3 | formula_repro |
Can combining English intermediate-task training with few-shot adaptation in the target language further improve zero-shot performance on no…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-10 03:36 |
| c2f07122-240e-4f… | Gate 3 | formula_repro |
How does the choice of intermediate task domain (e.g., natural language inference, sentiment analysis) affect cross-lingual zero-shot perfor…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-10 03:36 |
| 478950c1-47d0-45… | Gate 3 | formula_repro |
Does fine-tuning a pretrained multilingual model on English intermediate tasks improve zero-shot performance on non-English multimodal bench…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-07-10 03:36 |
| 8d980ea2-0f67-44… | Gate 3 | formula_repro |
How does the scaling of model size affect the efficacy of intermediate-task training for zero-shot cross-lingual transfer in XTREME-R?
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-10 03:36 |
| 684de976-bd27-45… | Gate 3 | formula_repro |
Can adversarial fine-tuning on intermediate English tasks improve the zero-shot cross-lingual transfer generalization of Bloom models, as me…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-10 03:36 |
| 8aaeef3e-12b1-4e… | Gate 3 | formula_repro |
How does varying the number of intermediate tasks in English intermediate-task training affect the zero-shot cross-lingual transfer performa…
COUNTEREXAMPLE HUNTER: 7.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.0/10 | 2026-07-10 03:36 |
| 8b75b109-ae6b-44… | Gate 3 | formula_repro |
How does the alignment of intermediate-task fine-tuning objectives (e.g., natural language inference vs. paraphrase detection) with target-t…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.2/10
|
6.1/10 | 2026-07-10 03:36 |
| 59641f13-782f-49… | Gate 3 | formula_repro |
How does the incorporation of domain-specific lexicons in zero-shot cross-lingual retrieval models trained on code-switched data affect perf…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-07-09 21:32 |
| f98f51bd-996a-43… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on code-switched data compare to models fine-tuned on parallel …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-09 21:32 |
| c94d82b1-812f-4c… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual transfer on XTREME-R compare when using multilingual intermediate tasks versus monolingu…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 1.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.1/10 | 2026-07-09 21:25 |
| 177db8b1-12e3-44… | Gate 3 | formula_repro |
How does the proportion of code-switched tokens in training data affect the cross-lingual retrieval performance of dense retrieval models on…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-07-09 15:21 |
| ac795d2e-8399-4a… | Gate 3 | formula_repro |
How does the token efficiency of cross-lingual retrieval models trained on code-switched data compare to standard zero-shot models when eval…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-09 15:17 |
| ec03fe07-25c6-4f… | Gate 3 | formula_repro |
Do multimodal models trained on artificially code-switched text-to-image pairs outperform text-only models in zero-shot cross-lingual image-…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 8.5/10
|
6.7/10 | 2026-07-09 15:15 |
| ce373b9d-6388-43… | Gate 3 | formula_repro |
What is the effect of combining English intermediate-task fine-tuning with multitask learning on target task performance in XTREME?
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-09 15:12 |
| 506ab53b-e424-4f… | Gate 3 | formula_repro |
How does the scaling of model size (e.g., 1B vs. 10B parameters) impact the effectiveness of domain-specific intermediate-task training for …
COUNTEREXAMPLE HUNTER: 7.0/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.0/10 | 2026-07-09 15:12 |
| 6af713c6-4956-45… | Gate 3 | formula_repro |
How does intermediate-task training on domain-specific English benchmarks (e.g., legal, medical) affect zero-shot cross-lingual transfer per…
COUNTEREXAMPLE HUNTER: 7.3/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.2/10
|
7.3/10 | 2026-07-09 15:10 |
| 5e995e21-f2a3-4b… | Gate 3 | formula_repro |
Does the hybrid batch training strategy improve zero-shot retrieval robustness to noisy code-switching queries in the BEIR benchmark compare…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.0/10
|
6.8/10 | 2026-07-09 09:09 |
| 5848c771-fb15-40… | Gate 3 | formula_repro |
How does the synergistic hybrid batch training strategy compare to standard domain-adversarial training in improving cross-lingual retrieval…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
6.3/10 | 2026-07-09 09:07 |
| d1523d0d-04a8-4c… | Gate 3 | formula_repro |
What is the impact of varying the proportion of monolingual, cross-lingual, and multilingual samples in hybrid batch training on inference t…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.5/10
|
6.2/10 | 2026-07-09 09:06 |
| f978e321-dc44-4f… | Gate 3 | formula_repro |
How does the scaling of intermediate-task training data size affect zero-shot cross-lingual performance on XTREME-R when comparing English-o…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-09 03:05 |
| 97b7c5f6-44fd-4f… | Gate 3 | formula_repro |
Does increasing the number of diverse high-resource non-English languages in multi-task intermediate training lead to diminishing returns in…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 8.5/10
|
6.7/10 | 2026-07-09 02:52 |
| 7c816cf9-e382-49… | Gate 3 | formula_repro |
Does the order of intermediate tasks (e.g., sequential vs. random) during English intermediate-task training affect the robustness of zero-s…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.4/10 | 2026-07-09 02:45 |
| fcacb1e4-51ce-4a… | Gate 3 | formula_repro |
How does the selection of intermediate English task difficulty impact the zero-shot cross-lingual transfer performance on XTREME-R benchmark…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-09 02:43 |
| d5cb2917-e4dc-45… | Gate 3 | formula_repro |
How does the performance of domain-specific intermediate-task training on legal or biomedical tasks compare to general English intermediate …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-07-08 20:41 |
| ebadefd7-ce3c-45… | Gate 3 | formula_repro |
What is the effect of multimodal intermediate task training on zero-shot cross-lingual performance compared to text-only intermediate tasks …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-07-08 20:39 |
| 45f4cba8-fe8d-49… | Gate 3 | formula_repro |
How does the choice of intermediate task domain impact zero-shot cross-lingual performance on the XTREME benchmark when using models larger …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-08 20:38 |
| 7d379fbb-4a60-4d… | Gate 3 | formula_repro |
How does the inference efficiency (measured in tokens/second and latency) of zero-shot cross-lingual retrieval models trained on code-switch…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-08 20:36 |
| bcdaa065-43f0-48… | Gate 3 | formula_repro |
What is the effect of model size (e.g., 1B vs. 10B parameters) on the performance gap between English and non-English intermediate tasks in …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-08 20:33 |
| 4724bbe9-b230-42… | Gate 2 | unknown | What is the impact of model size scaling (e.g., 7B vs. 30B parameters) on the effectiveness of English intermediate-task training for zero-s… | - | 2026-07-08 20:30 |
| 78c5715a-1570-42… | Gate 3 | formula_repro |
How does the performance gap between English and non-English target tasks in XTREME-R scale when using multilingual intermediate tasks (e.g.…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-08 20:30 |
| 4197a84c-341d-43… | Gate 3 | formula_repro |
What is the impact of scaling up the size of intermediate-task datasets on the zero-shot cross-lingual transfer performance of models on the…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.3/10 | 2026-07-08 20:25 |
| f06e54db-af1a-47… | Gate 3 | formula_repro |
How does the choice of intermediate-task training dataset size and domain relevance impact zero-shot cross-lingual transfer performance on X…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.2/10
|
5.5/10 | 2026-07-08 20:24 |
| 87c10cec-4d73-46… | Gate 3 | formula_repro |
How does adaptive batch sampling during sequential fine-tuning impact the zero-shot performance of XLM-R on euphemism detection in low-resou…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
6.2/10 | 2026-07-08 20:23 |
| 181fe369-b7d0-4c… | Gate 3 | formula_repro |
Does intermediate-task training on typologically diverse non-English languages improve few-shot transfer accuracy on XTREME-R classification…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.5/10
|
7.0/10 | 2026-07-08 20:23 |
| a57de557-74cf-46… | Gate 3 | formula_repro |
What is the impact of domain-specific intermediate tasks (e.g., legal, medical) on zero-shot cross-lingual transfer performance for non-Engl…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-08 14:20 |
| b517d425-2d39-4f… | Gate 3 | formula_repro |
To what extent does intermediate-task training on code generation tasks improve zero-shot cross-lingual transfer performance on XTREME-R com…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-08 14:19 |
| 5c5336df-89a8-4f… | Gate 3 | formula_repro |
Does the performance gain from intermediate-task training in English persist when evaluating on the XTREME-R benchmark, which includes retri…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-08 14:19 |
| 963137b5-817a-4e… | Gate 3 | formula_repro |
How does the performance of intermediate-task training on English compare to direct target-language fine-tuning when evaluated on the XGLUE …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-08 14:19 |
| 7233bd06-3f6d-4f… | Gate 3 | formula_repro |
Does scaling the number of English intermediate tasks in fine-tuning affect the adversarial robustness of zero-shot cross-lingual transfer o…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-08 14:19 |
| 03937d7f-568d-45… | Gate 3 | formula_repro |
Can intermediate-task training in English improve inference efficiency (latency/throughput) in zero-shot cross-lingual settings on XTREME-R …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-08 14:18 |
| 05cca66f-c0c1-49… | Gate 3 | formula_repro |
Does the order or sequence of intermediate language-understanding tasks (e.g., NLI before QA) impact zero-shot cross-lingual transfer perfor…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-08 14:18 |
| 9285a029-060e-4f… | Gate 3 | formula_repro |
How does intermediate-task training on low-resource language tasks (e.g., XTREME-R) compare to English intermediate-task training in zero-sh…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-08 14:18 |
| e8420509-72cf-49… | Gate 3 | formula_repro |
How does intermediate-task training with multimodal benchmarks like MMMU affect zero-shot cross-lingual transfer accuracy on text-based task…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-08 14:18 |
| 931dd776-0729-4c… | Gate 3 | formula_repro |
How does the scaling of intermediate-task training data size affect zero-shot cross-lingual transfer accuracy on XTREME-R, and what is the o…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-07-08 14:18 |
| 87937581-7492-49… | Gate 3 | formula_repro |
Does intermediate-task training in English improve multilingual alignment in models across different domains (e.g., legal, medical) as measu…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 6.5/10
|
7.6/10 | 2026-07-08 08:15 |
| f5f18118-8c91-4d… | Gate 3 | formula_repro |
How does the choice of intermediate task difficulty (easy vs. hard) influence zero-shot cross-lingual transfer performance in multilingual m…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-08 08:15 |
| 5f4c129d-5c44-46… | Gate 3 | formula_repro |
What is the impact of varying the number of intermediate language-understanding tasks on the zero-shot cross-lingual transfer capabilities o…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 6.5/10
|
8.1/10 | 2026-07-08 08:15 |
| 639cac0d-b0df-47… | Gate 3 | formula_repro |
How does intermediate-task training on multilingual data affect zero-shot performance on the XTREME benchmark compared to English-only inter…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-08 08:12 |
| fcb3b234-7fff-43… | Gate 3 | formula_repro |
How does the number of diverse intermediate language understanding tasks scale with the zero-shot cross-lingual transfer performance on XTRE…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-07-08 08:12 |
| 7a36578f-869c-43… | Gate 3 | formula_repro |
What is the impact of using different intermediate task sizes on the cross-lingual transfer capabilities of language models, measured by acc…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-08 08:10 |
| fe1711c5-b720-4d… | Gate 3 | formula_repro |
What is the effect of model size (e.g., comparing mT5-base vs. mT5-large) on the effectiveness of English intermediate-task training for zer…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.6/10 | 2026-07-08 08:10 |
| f0f5b9b8-753e-40… | Gate 3 | formula_repro |
How does the choice of intermediate task complexity (e.g., XNLI vs. TyDiQA) impact the zero-shot cross-lingual transfer performance on non-I…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-07-08 08:08 |
| d482b8c8-61cc-40… | Gate 2 | unknown | What is the impact of model size (e.g., XLM-R-base vs. XLM-R-large) on cross-lingual euphemism detection performance when using sequential f… | - | 2026-07-08 08:08 |
| 5013e4f7-275d-41… | Gate 3 | formula_repro |
Does intermediate-task training on non-English languages other than English yield comparable or better zero-shot cross-lingual performance c…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-08 08:07 |
| 7f46b379-016e-4e… | Gate 3 | formula_repro |
How does intermediate-task training influence the reasoning capabilities of multilingual models on cross-lingual arithmetic and logical reas…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-07-08 05:58 |
| 23e3aa02-ac4f-48… | Gate 3 | formula_repro |
How does intermediate-task training with domain-specific English corpora affect zero-shot cross-lingual performance on XTREME tasks compared…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.6/10 | 2026-07-08 05:58 |
| b52c3939-3b09-4b… | Gate 3 | formula_repro |
How does fine-tuning a multilingual model on a mix of English and non-English intermediate tasks affect zero-shot performance on XTREME comp…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-07-08 05:56 |
| 6a0611e3-395f-40… | Gate 3 | formula_repro |
Does intermediate-task training on code-related tasks (e.g., code generation, code reasoning) improve zero-shot cross-lingual performance on…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-07-08 05:56 |
| ebb6de3f-4114-44… | Gate 3 | formula_repro |
Does intermediate-task training with English-language tasks improve performance on XTREME-R for models with different architectures (e.g., t…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 6.5/10
|
8.2/10 | 2026-07-08 05:52 |
| 800542de-ce17-44… | Gate 3 | formula_repro |
How does scaling model size (e.g., 1B vs. 10B parameters) affect the robustness of zero-shot cross-lingual transfer on XTREME-R adversarial …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.5/10 | 2026-07-08 05:49 |
| 33daa49f-1c2a-49… | Gate 3 | formula_repro |
How does the choice of intermediate task language (non-English) affect zero-shot cross-lingual transfer accuracy in multilingual models like…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.1/10 | 2026-07-08 05:47 |
| c8332b4f-b6ad-4d… | Gate 3 | formula_repro |
Does scaling the number of non-English intermediate tasks (e.g., 3 vs. 9 tasks) improve zero-shot cross-lingual transfer on XTREME PAWS-X, a…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-08 05:45 |
| 11af2412-3a06-40… | Gate 3 | formula_repro |
To what extent does the order of intermediate-task fine-tuning (i.e., sequential vs. multi-task learning) impact zero-shot cross-lingual tra…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.4/10 | 2026-07-07 23:45 |
| 4105de18-6385-4b… | Gate 3 | formula_repro |
Does applying intermediate-task training on code-related tasks enhance the alignment of multilingual representations for zero-shot transfer …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-07 23:45 |
| ce0b9a39-da1c-4e… | Gate 3 | formula_repro |
Does intermediate-task training with reasoning tasks (e.g., GSM8K or MMLU) improve zero-shot cross-lingual syntactic generalization on XTREM…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-07-07 23:45 |
| 15b0adb1-69b7-46… | Gate 3 | formula_repro |
Do multimodal intermediate tasks (e.g., image captioning, V&L tasks) improve zero-shot cross-lingual transfer performance on XTREME tasks co…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-07-07 23:45 |
| 730f637d-51b1-4c… | Gate 3 | formula_repro |
How does scaling the number of English intermediate language understanding tasks affect the inference efficiency (e.g., latency, throughput)…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.4/10 | 2026-07-07 23:45 |
| 04fbbc5b-1a53-4f… | Gate 3 | formula_repro |
Does intermediate-task training on English adversarial datasets like ANLI improve zero-shot robustness scores on the XTREME benchmark for la…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.2/10 | 2026-07-07 17:45 |
| 518f7a24-eb4c-49… | Gate 3 | formula_repro |
How does the performance of intermediate-task training on English data compare to multilingual intermediate-task training when transferring …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-07 17:45 |
| def8fc46-d2b6-43… | Gate 3 | formula_repro |
Does intermediate-task training on multimodal reasoning datasets enhance zero-shot cross-lingual transfer accuracy on natural language infer…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
8.6/10 | 2026-07-07 17:44 |
| a2ffc254-86fe-49… | Gate 3 | formula_repro |
Does the effectiveness of English intermediate-task training on zero-shot cross-lingual transfer vary depending on the typological similarit…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-07 17:40 |
| c7881a79-1e60-4f… | Gate 3 | formula_repro |
How does the impact of intermediate-task training on zero-shot cross-lingual transfer performance compare between XTREME-R and XTREME when u…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.2/10 | 2026-07-07 17:40 |
| 83ee6131-343f-4d… | Gate 3 | formula_repro |
What is the impact of using multilingual intermediate tasks instead of English-only tasks on zero-shot cross-lingual transfer performance in…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-07-07 17:40 |
| 9b70a031-9a09-49… | Gate 3 | formula_repro |
Does intermediate-task training on English improve zero-shot cross-lingual performance on the XTREME-R benchmark compared to direct fine-tun…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-07 11:38 |
| 411590a4-1f00-46… | Gate 3 | formula_repro |
How does task diversity in English intermediate-task training affect zero-shot cross-lingual performance on XTREME-R, particularly for seman…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.2/10
|
5.5/10 | 2026-07-07 11:38 |
| a73c0c2f-5cbf-48… | Gate 3 | formula_repro |
How does intermediate-task training on domain-specific English datasets (e.g., legal, biomedical) affect cross-lingual transfer performance …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-07 11:37 |
| 7befddfc-4fb3-4e… | Gate 3 | formula_repro |
To what extent does intermediate-task training with English question-answering tasks generalize to cross-lingual retrieval tasks in XTREME-R…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-07 11:36 |
| 402ba3be-e898-41… | Gate 3 | formula_repro |
Does scaling the model size (from 100M to 10B+ parameters) affect the performance gap between intermediate-task training and direct fine-tun…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.3/10 · REPLICATION ATTACKER: 6.5/10
|
7.1/10 | 2026-07-07 11:34 |
| b921edba-e18b-4c… | Gate 3 | formula_repro |
Does scaling the size of the pretrained model impact the effectiveness of English intermediate-task training for zero-shot cross-lingual tra…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 3.0/10
|
6.6/10 | 2026-07-07 11:34 |
| fead3bd8-a3fa-48… | Gate 3 | formula_repro |
How does cross-lingual intermediate-task training (using multiple non-English languages) compare to English-only intermediate training in te…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-07 11:33 |
| 8fa5788e-4634-42… | Gate 2 | unknown | What is the impact of varying the number of intermediate languages in sequential fine-tuning on the zero-shot cross-lingual transfer accurac… | - | 2026-07-07 11:07 |
| 38ae3f3a-6fa4-48… | Gate 3 | formula_repro |
How does the effectiveness of English intermediate-task training on zero-shot cross-lingual transfer compare to training on intermediate tas…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-07-07 05:33 |
| 5ebfe0cd-0f58-4c… | Gate 3 | formula_repro |
How does intermediate-task training on non-English syntactic benchmarks compare to multilingual intermediate training in terms of zero-shot …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.6/10 | 2026-07-07 05:30 |
| 6c95af7e-3689-42… | Gate 3 | formula_repro |
Does the effectiveness of intermediate-task training for zero-shot cross-lingual transfer in XTREME vary with the linguistic complexity of t…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
5.6/10 | 2026-07-07 05:29 |
| 9dc4dc6b-f49f-48… | Gate 3 | formula_repro |
What is the impact of task diversity in intermediate training on the robustness of multilingual models in zero-shot cross-lingual transfer, …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
8.6/10 | 2026-07-07 05:27 |
| 4bf74eec-b439-4b… | Gate 2 | unknown | How does the inclusion of intermediate language fine-tuning (e.g., from English → Spanish → Turkish) compare to direct sequential fine-tunin… | - | 2026-07-07 02:30 |
| 406f6d9e-ff82-4e… | Gate 3 | formula_repro |
What is the impact of using non-English intermediate tasks (e.g., from XNLI or TyDiQA) on zero-shot cross-lingual transfer performance compa…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 6.5/10
|
6.1/10 | 2026-07-06 23:26 |
| fa78a2d2-49e1-43… | Gate 3 | formula_repro |
Does the effectiveness of English intermediate-task training on cross-lingual transfer vary across different language families in the XTREME…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.0/10 | 2026-07-06 23:26 |
| 82aec9df-5413-4d… | Gate 3 | formula_repro |
To what extent does domain-specific intermediate-task training (e.g., legal, medical) improve zero-shot transfer performance on XTREME for l…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.8/10 | 2026-07-06 23:26 |
| ebac71db-e7f5-4d… | Gate 3 | formula_repro |
What is the effect of scaling the number of languages in intermediate-task training (e.g., multilingual vs. English-only) on zero-shot cross…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
4.7/10 | 2026-07-06 23:26 |
| ecce3153-d200-44… | Gate 3 | formula_repro |
What is the impact of combining intermediate-task training with cross-lingual data augmentation on XTREME-R performance, compared to using e…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-06 17:17 |
| 36d3ebfe-1234-40… | Gate 3 | formula_repro |
Does the order of intermediate tasks in multi-task training influence zero-shot cross-lingual transfer performance on XTREME-R classificatio…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-06 17:17 |
| 58ca9097-4fd4-47… | Gate 3 | formula_repro |
How does multi-task intermediate training with non-English tasks affect zero-shot cross-lingual performance on XTREME-R compared to English-…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-06 17:16 |
| 2d74fcd3-a755-49… | Gate 3 | formula_repro |
How does the scaling of model size influence the effectiveness of intermediate-task fine-tuning for zero-shot cross-lingual transfer, partic…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-06 17:16 |
| 353e83f6-40b2-4a… | Gate 3 | formula_repro |
How does the order of intermediate-task fine-tuning in multilingual models affect the robustness of zero-shot cross-lingual performance when…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-07-06 17:16 |
| 341930a5-2982-45… | Gate 3 | formula_repro |
How does multi-task intermediate fine-tuning on diverse non-English languages compare to English-only intermediate fine-tuning for zero-shot…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-07-06 17:16 |
| f0a2cebb-4ff9-47… | Gate 3 | formula_repro |
How does scaling the diversity of intermediate tasks (e.g., mixing English and multilingual tasks) impact zero-shot cross-lingual transfer a…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-07-06 17:15 |
| a3b69e76-aa82-4b… | Gate 3 | formula_repro |
How does the hybrid batch training method compare to adversarial training in improving zero-shot cross-lingual retrieval accuracy on legal d…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.0/10 | 2026-07-06 11:13 |
| 97ec1a75-7710-4f… | Gate 3 | formula_repro |
Do multilingual intermediate tasks outperform English-only intermediate tasks for zero-shot cross-lingual transfer across diverse typologies…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.7/10 | 2026-07-06 11:12 |
| 644bb2dd-f15e-46… | Gate 3 | formula_repro |
How does the scaling of model size beyond 750M parameters affect the zero-shot cross-lingual transfer performance on XTREME-R when using dif…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-07-06 11:12 |
| dd2b3a29-c88b-4f… | Gate 3 | formula_repro |
What is the impact of model size scaling on the effectiveness of English intermediate-task training for zero-shot cross-lingual transfer, as…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-06 11:12 |
| 24c3bb8c-d457-48… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual transfer in XTREME benchmark compare between models fine-tuned on domain-specific interm…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-07-06 11:12 |
| 836391f0-85d2-48… | Gate 3 | formula_repro |
Does scaling the number of English intermediate tasks yield diminishing returns for zero-shot cross-lingual performance on the XNLI benchmar…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-07-06 11:12 |
| 692a7635-19bc-47… | Gate 3 | formula_repro |
How does intermediate-task training on typologically diverse source languages affect zero-shot accuracy on low-resource languages within the…
COUNTEREXAMPLE HUNTER: 8.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
8.9/10 | 2026-07-06 11:12 |
| 3c02c7e4-dee1-40… | Gate 2 | unknown | How does the integration of synthetic code-switched training data impact the zero-shot cross-lingual retrieval performance (mAP) of language… | - | 2026-07-06 08:47 |
| 9518af70-a756-47… | Gate 2 | unknown | How does the performance gap in F1 score between sequential and simultaneous multilingual fine-tuning of XLM-R vary when applied to low-reso… | - | 2026-07-06 08:45 |
| 4b67a093-bc17-43… | Gate 2 | unknown | What is the impact of typologically diverse language pairings (e.g., English-Turkish vs. Yoruba-Chinese) on the F1 score performance of XLM-… | - | 2026-07-06 08:44 |
| acf5dca9-3900-4c… | Gate 2 | unknown | How does the order of language exposure in sequential fine-tuning affect the robustness of XLM-R and mBERT against StressGAN-generated adver… | - | 2026-07-06 08:44 |
| 4b0df5f0-db98-43… | Gate 2 | unknown | What is the impact of typological similarity between language pairs on cross-lingual transfer performance in euphemism detection, as measure… | - | 2026-07-06 08:44 |
| e7088f7f-959d-48… | Gate 3 | formula_repro |
What is the impact of intermediate-task training on the robustness of zero-shot cross-lingual transfer when evaluating on adversarial or noi…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.2/10 | 2026-07-06 05:11 |
| fc1711a8-dc47-48… | Gate 3 | formula_repro |
What is the effect of English intermediate-task fine-tuning on the transfer performance of XLM-R for cross-lingual natural language inferenc…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-07-06 05:11 |
| f8c356fd-66e5-40… | Gate 3 | formula_repro |
To what extent does the order of intermediate-task training (sequential vs. interleaved) affect the zero-shot cross-lingual transfer perform…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 9.2/10
|
6.2/10 | 2026-07-06 05:11 |
| cbe3c606-0594-4b… | Gate 3 | formula_repro |
How does the performance of XLM-R on zero-shot cross-lingual logical inference tasks in XTREME-R compare to mT5 when both models undergo Eng…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 2.0/10
|
5.3/10 | 2026-07-06 05:09 |
| 55378023-1709-42… | Gate 3 | formula_repro |
How does English intermediate-task training impact the robustness of XLM-R Large against typological divergence in zero-shot cross-lingual l…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-06 05:09 |
| 01bedcd8-5b1c-49… | Gate 3 | formula_repro |
Does the efficacy of English intermediate-task training for zero-shot cross-lingual logical inference on XLM-R degrade as model scale increa…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.5/10
|
6.8/10 | 2026-07-06 05:09 |
| a855af96-b322-4a… | Gate 3 | formula_repro |
How does scaling pretrained model size from base to large variants impact the retention of cross-lingual alignment when fine-tuned on divers…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-07-05 23:07 |
| 53edb8d6-bf32-46… | Gate 3 | formula_repro |
How does the performance gap between multilingual and English-only intermediate-task training on XTREME vary across different parameter scal…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-05 23:07 |
| e04654a8-bb43-43… | Gate 3 | formula_repro |
Does English intermediate-task training degrade zero-shot transfer performance on typologically distant languages in XTREME compared to ling…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.5/10 | 2026-07-05 23:07 |
| 971c98ef-1fb7-4f… | Gate 3 | formula_repro |
Does intermediate-task training on English data improve robustness against typological divergence in zero-shot cross-lingual transfer for la…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.8/10 | 2026-07-05 23:07 |
| b11247d6-3d0b-4e… | Gate 3 | formula_repro |
How robust is English intermediate-task training to domain shifts when evaluated on low-resource languages within the XTREME natural languag…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-07-05 23:06 |
| be8a954b-c5b8-46… | Gate 3 | formula_repro |
How does intermediate-task training on non-English source languages affect zero-shot transfer accuracy on XCOPA compared to English-only int…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.3/10 · REPLICATION ATTACKER: 7.5/10
|
7.1/10 | 2026-07-05 23:06 |
| ccdbdcf9-5d04-4f… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on the XTREME benchmark diminish when applied to massively multilingual mo…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-05 23:04 |
| 527737c7-4f29-42… | Gate 2 | unknown | To what extent does fine-tuning multilingual embeddings on domain-specific artificially code-switched data improve zero-shot cross-lingual r… | - | 2026-07-05 18:18 |
| 380cb179-688f-46… | Gate 3 | formula_repro |
How does intermediate-task training on English entailment tasks affect the robustness of zero-shot cross-lingual transfer on XNLI under adve…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-07-05 17:04 |
| d85c5f4b-e41f-46… | Gate 3 | formula_repro |
What is the impact of scaling the size of the pretrained models on the effectiveness of English intermediate-task training for zero-shot cro…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-07-05 17:04 |
| 06bcb851-5ea9-4d… | Gate 3 | formula_repro |
What is the impact of varying the number or diversity of intermediate English tasks on zero-shot cross-lingual transfer performance in multi…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.8/10 | 2026-07-05 17:04 |
| d6cd9aba-3c28-41… | Gate 3 | formula_repro |
How does the zero-shot cross-lingual transfer performance of multilingual models compare to monolingual models when trained on English inter…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.5/10
|
7.0/10 | 2026-07-05 17:04 |
| b757dfc2-fbc2-4c… | Gate 3 | formula_repro |
How does intermediate-task training on non-English source languages impact zero-shot transfer performance to typologically distant targets o…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.3/10 | 2026-07-05 17:03 |
| 31c17b61-748f-41… | Gate 3 | formula_repro |
How does semantic similarity between multilingual intermediate tasks and target tasks in XTREME-R correlate with zero-shot cross-lingual tra…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 6.5/10
|
7.4/10 | 2026-07-05 17:02 |
| 6062b432-3d6b-40… | Gate 3 | formula_repro |
What is the impact of varying the proportion of code-switched queries in training data on the robustness of zero-shot cross-lingual retrieva…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.2/10 | 2026-07-05 11:02 |
| dac0b4d5-3d8d-41… | Gate 3 | formula_repro |
How does adversarial training on the intermediate task affect the zero-shot cross-lingual transfer performance on XTREME compared to standar…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 7.5/10
|
8.2/10 | 2026-07-05 11:02 |
| 9bf107ed-c391-48… | Gate 3 | formula_repro |
What is the effect of domain-specific intermediate-task training (e.g., legal or biomedical) on zero-shot cross-lingual transfer performance…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.2/10
|
6.1/10 | 2026-07-05 11:00 |
| 3d3cfe0f-6835-4e… | Gate 3 | formula_repro |
What is the impact of integrating pre-trained multilingual embeddings with teacher-student learning on the accuracy and robustness of cross-…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 8.0/10
|
7.3/10 | 2026-07-05 05:00 |
| 9bda959c-e4f9-4e… | Gate 3 | formula_repro |
How does the performance of cross-lingual NER models trained via annotation projection compare to that of multilingual language models like …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 2.5/10
|
6.5/10 | 2026-07-05 04:55 |
| 6c88c86d-932e-42… | Gate 3 | formula_repro |
Does English intermediate-task training improve robustness against adversarial perturbations in non-English languages when evaluated on the …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-04 22:55 |
| 09e19d49-354f-42… | Gate 3 | formula_repro |
How does the accuracy of projection-based cross-lingual NER scale with increasing ratios of image-text parallel data compared to text-only b…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.5/10 | 2026-07-04 22:55 |
| 2ec27aca-4798-4b… | Gate 3 | formula_repro |
To what extent does the choice of alignment metric (e.g., Wasserstein distance, KL divergence) in adversarial cross-lingual NER models affec…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 4.5/10
|
7.5/10 | 2026-07-04 22:53 |
| 132a9aa2-daa8-42… | Gate 3 | formula_repro |
What is the impact of multilingual intermediate-task training (using multiple source languages) versus monolingual English intermediate-task…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-04 16:51 |
| 8f450884-681e-44… | Gate 3 | formula_repro |
How does multilingual intermediate-task training (vs. English-only) affect zero-shot cross-lingual transfer performance for low-resource lan…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.5/10 | 2026-07-04 16:50 |
| 8d9c0600-8114-42… | Gate 3 | formula_repro |
Does scaling the size of the pretrained model amplify the benefits of English intermediate-task training for zero-shot transfer on low-resou…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-07-04 10:48 |
| d67fa29a-4968-45… | Gate 3 | formula_repro |
Does the hybrid batch training approach improve monolingual retrieval metrics on BEIR while maintaining cross-lingual zero-shot accuracy on …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-07-04 04:43 |
| cd30c9eb-b95e-47… | Gate 3 | formula_repro |
How does the choice of intermediate task granularity (e.g., fine-grained vs. coarse-grained) impact zero-shot cross-lingual transfer perform…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.8/10 | 2026-07-03 22:41 |
| 41e91c47-e63e-4e… | Gate 3 | formula_repro |
Does the effectiveness of English intermediate-task training for zero-shot cross-lingual transfer hold consistently across different languag…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-07-03 22:41 |
| db6b3f84-b398-49… | Gate 3 | formula_repro |
Does increasing the parameter count of multilingual pretrained models reduce the performance gap between English intermediate-task training …
COUNTEREXAMPLE HUNTER: 7.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.0/10 | 2026-07-03 22:41 |
| a289f19e-10eb-40… | Gate 3 | formula_repro |
What is the impact of model size scaling (e.g., 1B, 3B, 10B parameters) on the trade-off between monolingual and cross-lingual retrieval per…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 0.5/10
|
5.5/10 | 2026-07-03 22:33 |
| fe811201-ce08-49… | Gate 3 | formula_repro |
Does the effectiveness of English intermediate-task training on zero-shot cross-lingual transfer vary across different language families in …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.4/10 | 2026-07-03 16:30 |
| de98f41c-66f8-4f… | Gate 3 | formula_repro |
To what extent does English intermediate-task training improve zero-shot transfer performance on low-resource languages within the XTREME be…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 4.3/10 · REPLICATION ATTACKER: 7.5/10
|
6.4/10 | 2026-07-03 16:30 |
| 60b8630d-3de5-45… | Gate 3 | formula_repro |
To what extent does English intermediate-task training improve robustness against typological distance in zero-shot cross-lingual transfer o…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 8.5/10
|
6.7/10 | 2026-07-03 16:30 |
| bc83ed60-2a7a-4c… | Gate 3 | formula_repro |
Does the effectiveness of intermediate-task training for zero-shot cross-lingual transfer vary significantly when using intermediate tasks i…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.0/10 | 2026-07-03 10:29 |
| d3626aad-595f-4d… | Gate 3 | formula_repro |
How does English intermediate-task training impact the robustness of zero-shot cross-lingual transfer on XTREME when evaluated against adver…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-07-03 10:29 |
| cad742bb-8637-42… | Gate 3 | formula_repro |
Does scaling the volume of artificially code-switched training data yield diminishing returns for zero-shot cross-lingual retrieval performa…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.5/10 | 2026-07-03 10:29 |
| 3835fb25-aba0-48… | Gate 3 | formula_repro |
How does the scaling behavior of zero-shot cross-lingual retrieval models trained on artificially code-switched data differ from models fine…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 1.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.4/10 | 2026-07-03 10:29 |
| 5f3313e4-48a2-4a… | Gate 3 | formula_repro |
How does the robustness of zero-shot cross-lingual retrieval models trained on code-switched data vary across different language pairs (e.g.…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-07-03 04:29 |
| 8c7c87f1-82be-43… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data scale with the size of the b…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 2.0/10
|
6.0/10 | 2026-07-03 04:29 |
| ddaf7d86-6d9b-4e… | Gate 3 | formula_repro |
How well do zero-shot cross-lingual retrieval models trained on artificially code-switched data generalize to low-resource languages not pre…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-03 04:29 |
| aa368bcc-fc6c-44… | Gate 3 | formula_repro |
To what extent does scaling the model size (e.g., base vs. large vs. extra-large) influence the effectiveness of English intermediate-task t…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.8/10 | 2026-07-03 04:29 |
| 265a876c-3db8-4c… | Gate 3 | formula_repro |
What is the impact of domain adaptation techniques (e.g., domain-specific fine-tuning) on the robustness of English intermediate-task-tuned …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.8/10 | 2026-07-03 04:29 |
| 55206d0d-fc16-45… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual transfer on XTREME-R tasks compare when using different combinations of English intermed…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-03 04:29 |
| 8014e20a-13b9-4e… | Gate 3 | formula_repro |
Does the choice of intermediate task (e.g., natural language inference vs. question answering) affect the degree of performance improvement …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.4/10 | 2026-07-03 04:29 |
| 988ef48b-1605-47… | Gate 3 | formula_repro |
What is the impact of model size on the effectiveness of English intermediate-task training for zero-shot cross-lingual transfer, comparing …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-03 04:29 |
| 25c0a4e8-6737-46… | Gate 3 | formula_repro |
How does the zero-shot cross-lingual performance of models trained on English intermediate tasks compare to multilingual intermediate-task t…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-03 04:29 |
| f51e721b-7196-45… | Gate 3 | formula_repro |
What is the robustness of zero-shot cross-lingual retrieval models trained on code-switched data when applied to documents with varying leve…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-02 22:29 |
| 8c3bece3-a45a-49… | Gate 3 | formula_repro |
Does increasing the volume of high-resource speech data during multilingual pretraining improve the robustness of SLAM-ASR models against no…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-02 22:29 |
| 3a74baf3-dc9c-40… | Gate 3 | formula_repro |
To what extent does fine-tuning zero-shot cross-lingual retrieval models on domain-specific code-switched data (e.g., legal or medical corpo…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-02 22:28 |
| 1a6a6b22-4bbe-4f… | Gate 3 | formula_repro |
Does the effectiveness of zero-shot cross-lingual retrieval models trained on code-switched data degrade when applied to low-resource langua…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.8/10 | 2026-07-02 22:28 |
| 1bd1fdd1-d70d-41… | Gate 3 | formula_repro |
How effective is adversarial training with artificially code-switched data in improving zero-shot cross-lingual retrieval performance on the…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 3.5/10
|
6.5/10 | 2026-07-02 22:28 |
| daf802ce-71bc-4f… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare to models fine-tuned…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-07-02 16:28 |
| 3309ef97-4dd4-4f… | Gate 3 | formula_repro |
How does the integration of adversarial training with artificially code-switched data improve the robustness of zero-shot cross-lingual retr…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.8/10 | 2026-07-02 16:21 |
| 27f76be0-f074-46… | Gate 3 | formula_repro |
How does the variation in the size of bilingual lexicons used for artificial code-switching affect the performance of multilingual language …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.2/10
|
6.4/10 | 2026-07-02 16:17 |
| f7039c85-8031-48… | Gate 3 | formula_repro |
What is the impact of domain-specific code-switched training data on zero-shot cross-lingual retrieval performance, measured by MRR on domai…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-07-02 15:59 |
| 50c8ae16-0789-4c… | Gate 3 | formula_repro |
How does varying the proportion of code-switched tokens in artificially generated training data affect zero-shot cross-lingual retrieval per…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.0/10 | 2026-07-02 15:59 |
| aa6db607-4e30-4e… | Gate 2 | unknown | How does the hybrid batch training strategy proposed in this paper compare to other multilingual pre-training techniques (e.g., mBERT, XLM-R… | - | 2026-07-02 13:56 |
| ba523cf3-ff96-4d… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on non-English targets in the XTREME benchmark diminish when the base mode…
COUNTEREXAMPLE HUNTER: 7.3/10 · CITATION AUDITOR: 2.3/10 · REPLICATION ATTACKER: 6.5/10
|
5.4/10 | 2026-07-02 09:59 |
| 08fbaebe-3e62-46… | Gate 3 | formula_repro |
Does scaling the number of diverse English intermediate tasks yield diminishing returns in cross-lingual generalization on XTREME-R compared…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
4.7/10 | 2026-07-02 09:59 |
| 323ba72b-4509-40… | Gate 3 | formula_repro |
Does English intermediate-task training degrade zero-shot cross-lingual transfer performance on low-resource languages within the XTREME ben…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-02 09:59 |
| dc1f8a23-af8d-40… | Gate 3 | formula_repro |
How does the effectiveness of English intermediate-task training vary across different language families in XTREME when controlling for mode…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.2/10
|
6.1/10 | 2026-07-02 09:53 |
| 60000bc3-6604-4f… | Gate 3 | formula_repro |
Does incorporating multilingual intermediate tasks enhance the reasoning capabilities of pretrained models in zero-shot cross-lingual settin…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.5/10 | 2026-07-02 09:48 |
| 4a57b572-6743-4f… | Gate 3 | formula_repro |
How does intermediate-task training on diverse non-English datasets affect zero-shot transfer performance on XTREME-R compared to English-on…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-07-02 09:48 |
| 822f0829-d336-49… | Gate 3 | formula_repro |
How does the scaling behavior of English intermediate-task training for zero-shot cross-lingual transfer vary across different model archite…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 4.5/10
|
6.2/10 | 2026-07-02 03:48 |
| 8ff4c836-efce-4b… | Gate 3 | formula_repro |
Does scaling the size of the pretrained multilingual model mitigate the performance gap between English-only and multilingual intermediate-t…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 1.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.2/10 | 2026-07-02 03:48 |
| 0a2cd3fc-9ca1-4d… | Gate 3 | formula_repro |
To what extent does adversarial fine-tuning on intermediate tasks improve model performance on non-English target tasks in XTREME-R compared…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-02 03:48 |
| 75970e40-a6e4-4c… | Gate 3 | formula_repro |
What is the impact of varying the size of the pretrained model on the robustness of intermediate-task training for zero-shot cross-lingual t…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-02 03:48 |
| ba7a9361-e985-4a… | Gate 3 | formula_repro |
Do multilingual intermediate-task training strategies outperform English-only intermediate-task training when evaluated on XTREME-R for zero…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-07-02 03:48 |
| e6c7368e-ebcf-41… | Gate 3 | formula_repro |
What is the impact of combining multiple non-English intermediate tasks from diverse languages on zero-shot cross-lingual transfer performan…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-07-01 21:46 |
| 0ca47119-204a-4b… | Gate 3 | formula_repro |
How does the choice of intermediate task (e.g., sentiment analysis, natural language inference) influence the robustness of zero-shot cross-…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-07-01 21:46 |
| 9eacad34-9e39-46… | Gate 3 | formula_repro |
What is the effect of scaling the number of target languages in zero-shot cross-lingual transfer on the performance gap between English and …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-07-01 21:45 |
| 830cec29-1e4e-44… | Gate 3 | formula_repro |
Does English intermediate-task training improve robustness against adversarial perturbations in zero-shot cross-lingual settings compared to…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
8.2/10 | 2026-07-01 21:43 |
| e5902724-9a7b-4d… | Gate 3 | formula_repro |
How does increasing the parameter count of English intermediate-task trained models impact the zero-shot cross-lingual performance gap on XT…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-07-01 15:39 |
| d3718cc7-710f-49… | Gate 3 | formula_repro |
How does the performance of intermediate-task transfer compare when using multilingual intermediate tasks versus English-only tasks on zero-…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.0/10 · REPLICATION ATTACKER: 7.5/10
|
6.0/10 | 2026-07-01 15:39 |
| 355737d8-3fce-4e… | Gate 3 | formula_repro |
How does intermediate-task training on non-English datasets compare to English intermediate tasks for zero-shot transfer performance on the …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-07-01 15:39 |
| 45946c68-2bd2-44… | Gate 3 | formula_repro |
Does the choice of intermediate-task linguistic complexity (e.g., semantic vs. syntactic tasks) influence the zero-shot cross-lingual transf…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.2/10
|
4.6/10 | 2026-07-01 15:39 |
| 29387235-64f1-4f… | Gate 3 | formula_repro |
What is the impact of domain-specific intermediate tasks (e.g., legal or medical) on zero-shot cross-lingual transfer performance in XNLI co…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-07-01 09:35 |
| bf1cafcf-73e6-4f… | Gate 3 | formula_repro |
Does the benefit of English intermediate-task training on zero-shot cross-lingual transfer persist when the target tasks in XTREME require c…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 6.5/10
|
6.1/10 | 2026-07-01 09:34 |
| a06fc2bb-50b5-46… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on zero-shot cross-lingual transfer scale with the size of the pretrained …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-07-01 03:22 |
| 1906d541-329e-42… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on XTREME-R tasks diminish when scaling to larger multilingual pretrained …
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-07-01 03:22 |
| 54496c53-db81-4c… | Gate 3 | formula_repro |
How does the accuracy on XTREME zero-shot transfer compare when intermediate-task training uses code-switched multilingual data versus monol…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-07-01 03:22 |
| d8fa094f-944f-4d… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on XTREME classification tasks diminish when scaling from base to large pr…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.2/10 · REPLICATION ATTACKER: 8.5/10
|
7.4/10 | 2026-07-01 03:22 |
| b31bb4c5-ac4f-49… | Gate 3 | formula_repro |
Does intermediate-task training on typologically diverse non-English datasets outperform English-only intermediate training for zero-shot tr…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.8/10 | 2026-07-01 03:20 |
| 80293841-85eb-48… | Gate 3 | formula_repro |
What is the impact of scaling the number of intermediate languages during training on the robustness of zero-shot cross-lingual transfer per…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-06-30 21:20 |
| e9a06320-59f9-47… | Gate 3 | formula_repro |
Does intermediate-task training on English adversarial datasets improve zero-shot adversarial robustness for low-resource languages in the X…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.3/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-06-30 21:19 |
| 7f257464-279a-49… | Gate 3 | formula_repro |
Does scaling the number of English intermediate training tasks yield diminishing returns for zero-shot performance on low-resource languages…
COUNTEREXAMPLE HUNTER: 7.3/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 6.5/10
|
6.0/10 | 2026-06-30 21:19 |
| 8f853428-9d67-43… | Gate 3 | formula_repro |
How does scaling the diversity of English intermediate tasks affect zero-shot cross-lingual accuracy on sentence retrieval within the XTREME…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 3.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.6/10 | 2026-06-30 21:19 |
| 3f7a5da1-5bfd-47… | Gate 3 | formula_repro |
Does English intermediate-task training improve robustness against typological distance in zero-shot cross-lingual transfer for language und…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 6.5/10
|
5.7/10 | 2026-06-30 21:19 |
| adf29182-b8df-4f… | Gate 3 | formula_repro |
How does the scaling of model parameters from base to extra-large affect the zero-shot cross-lingual transfer accuracy on XTREME when applyi…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.0/10 | 2026-06-30 21:19 |
| b09626ef-4dcd-48… | Gate 3 | formula_repro |
Does scaling the number of English intermediate tasks improve zero-shot transfer performance on non-English reasoning benchmarks within XTRE…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.2/10
|
5.7/10 | 2026-06-30 21:17 |
| 7c0105f7-f74e-45… | Gate 3 | formula_repro |
Does the performance gap between intermediate-task training and direct fine-tuning persist when scaling up the model size in zero-shot cross…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.2/10 | 2026-06-30 15:15 |
| 148d6fd6-790e-49… | Gate 3 | formula_repro |
What is the effect of using multilingual intermediate tasks (e.g., combining English and non-English tasks) on zero-shot cross-lingual trans…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.2/10
|
6.4/10 | 2026-06-30 15:15 |
| 6ce8955e-dc37-4e… | Gate 3 | formula_repro |
Does intermediate-task training on code-switched datasets improve zero-shot cross-lingual transfer accuracy on XTREME classification tasks c…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-06-30 15:13 |
| 2a059e4e-41a9-48… | Gate 3 | formula_repro |
How does intermediate-task fine-tuning on English datasets affect the robustness of zero-shot cross-lingual transfer under domain shift cond…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.1/10 | 2026-06-30 09:02 |
| fb6ea4ef-cfa0-4f… | Gate 3 | formula_repro |
Does adversarial training on English intermediate tasks improve robustness against synthetic noise in zero-shot cross-lingual transfer on th…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-06-30 09:02 |
| 27c3c93c-206b-4d… | Gate 3 | formula_repro |
How does intermediate-task training on English legal and biomedical datasets affect zero-shot cross-lingual transfer accuracy for low-resour…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 6.5/10
|
5.7/10 | 2026-06-30 08:58 |
| 3176076f-794f-41… | Gate 3 | formula_repro |
Does English intermediate-task training improve robustness against typological divergence in low-resource languages within the XTREME benchm…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-06-30 08:56 |
| dc2ca341-8965-45… | Gate 2 | unknown | How does the number of intermediate tasks in sequential fine-tuning affect zero-shot cross-lingual transfer accuracy on the XTREME-R benchma… | - | 2026-06-30 07:56 |
| 9b70a6ce-6ad1-48… | Gate 3 | formula_repro |
What is the impact of English intermediate-task difficulty on the robustness of zero-shot cross-lingual transfer to typologically distant la…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-06-30 04:09 |
| 93ca8478-ef77-4d… | Gate 3 | formula_repro |
How does intermediate-task training in non-English source languages with varying typological similarities (e.g., Hungarian vs. Russian) affe…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
4.7/10 | 2026-06-30 04:09 |
| e2a65a24-14c3-49… | Gate 3 | formula_repro |
Does intermediate-task training on multilingual syntactic parsing improve zero-shot transfer accuracy on XTREME-R compared to English-only i…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-06-30 04:09 |
| c422fcfb-e7b2-4c… | Gate 3 | formula_repro |
What is the impact of intermediate-task training on the robustness of zero-shot cross-lingual transfer to low-resource languages in the XTRE…
COUNTEREXAMPLE HUNTER: 2.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-06-30 04:02 |
| 635a9fdf-7cad-4d… | Gate 3 | formula_repro |
How does the scale of the pretrained model (e.g., BLOOM-176B vs. mT5-XXL) influence the magnitude of zero-shot cross-lingual transfer gains …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-06-29 21:44 |
| 1162fe8c-c983-48… | Gate 3 | formula_repro |
Does English intermediate-task training degrade robustness to typological divergence when evaluated on low-resource language subsets of the …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.8/10 | 2026-06-29 21:44 |
| 84862043-3a9d-47… | Gate 3 | formula_repro |
How does intermediate-task training on multimodal datasets affect zero-shot cross-lingual transfer performance on XTREME-R compared to text-…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.2/10
|
6.3/10 | 2026-06-29 21:40 |
| 33e69294-93fc-4b… | Gate 3 | formula_repro |
How does the scaling of intermediate task diversity (e.g., 5 vs. 15 tasks) impact zero-shot cross-lingual transfer performance on XTREME-R f…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.0/10
|
6.0/10 | 2026-06-29 15:40 |
| ae693872-097f-47… | Gate 3 | formula_repro |
Does the order of intermediate-task fine-tuning sequences influence the final zero-shot cross-lingual performance metrics on the XTREME benc…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.8/10 | 2026-06-29 15:40 |
| e32dc6eb-437b-40… | Gate 3 | formula_repro |
To what extent does English intermediate-task training improve robustness against typological distance when evaluating zero-shot transfer on…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-06-29 15:40 |
| 91c10821-4e91-4e… | Gate 3 | formula_repro |
How does the semantic diversity of intermediate tasks affect zero-shot cross-lingual transfer accuracy on the XTREME benchmark compared to s…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-06-29 15:40 |
| f3e38758-391a-45… | Gate 3 | formula_repro |
Does English intermediate-task training improve zero-shot cross-lingual transfer robustness on typologically distant languages in XTREME com…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-29 15:39 |
| 4c5389e1-e582-4b… | Gate 3 | formula_repro |
What is the impact of intermediate-task training on the robustness of cross-lingual transfer for low-resource languages within the XTREME be…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
8.6/10 | 2026-06-29 15:39 |
| bee34aef-f7aa-40… | Gate 3 | formula_repro |
How does the performance of non-English intermediate-task training compare to English intermediate-task training for zero-shot cross-lingual…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-29 15:36 |
| 17034fd1-c6a8-49… | Gate 3 | formula_repro |
Does intermediate-task training on English tasks improve zero-shot cross-lingual reasoning performance in multilingual models, as measured b…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-29 15:36 |
| 49c922f7-ee15-4d… | Gate 3 | formula_repro |
To what extent does English intermediate-task fine-tuning improve robustness against typographical noise in zero-shot cross-lingual settings…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-06-29 15:36 |
| 9623ffe5-b951-45… | Gate 3 | formula_repro |
Can the hybrid batch training approach be scaled to improve simultaneous zero-shot cross-lingual and multilingual retrieval on the TyDi QA b…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-06-29 09:35 |
| 2a5d9a58-9436-46… | Gate 3 | formula_repro |
How does intermediate-task training on multilingual pretraining models like XLM-R compare to English-only pretraining in terms of zero-shot …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-06-29 09:35 |
| cd03850f-9d60-4f… | Gate 3 | formula_repro |
How does domain-specific intermediate-task training (e.g., legal or medical tasks) affect zero-shot cross-lingual transfer performance compa…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.2/10
|
6.0/10 | 2026-06-29 09:35 |
| 7b6c77b6-8325-44… | Gate 3 | formula_repro |
Do multimodal intermediate tasks (e.g., image-text alignment) provide additional benefits over text-only intermediate tasks for zero-shot cr…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.3/10
|
6.1/10 | 2026-06-29 09:35 |
| 9a3cef16-b42d-4f… | Gate 3 | formula_repro |
Does the hybrid batch training strategy maintain robustness in zero-shot cross-lingual retrieval on MLDR when evaluated against domain-shift…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 2.0/10
|
5.7/10 | 2026-06-29 03:35 |
| 23b4506e-7120-40… | Gate 3 | formula_repro |
How does the hybrid batch training strategy impact zero-shot cross-lingual retrieval accuracy on the BUCC benchmark for low-resource languag…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-06-29 03:35 |
| 04209bd1-64f1-44… | Gate 3 | formula_repro |
Does the effectiveness of English intermediate-task training for zero-shot cross-lingual transfer scale with model size, as measured by accu…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.5/10 | 2026-06-29 03:35 |
| cae5b5dd-973f-47… | Gate 3 | formula_repro |
Does scaling the diversity of intermediate tasks in non-English high-resource languages (e.g., Spanish, French) improve zero-shot cross-ling…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.8/10 | 2026-06-29 03:33 |
| a9878558-b5c6-4d… | Gate 3 | formula_repro |
What is the effect of model size on the zero-shot cross-lingual transfer performance of multilingual models when fine-tuned on a diverse set…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-06-29 03:33 |
| a18feabc-7a33-4b… | Gate 3 | formula_repro |
Does increasing the diversity of non-English intermediate tasks enhance zero-shot transfer performance on low-resource languages within the …
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.1/10 | 2026-06-29 03:33 |
| 9db96c7b-5160-42… | Gate 3 | formula_repro |
Does fine-tuning a multilingual model on English intermediate tasks affect the robustness of zero-shot transfer to low-resource languages in…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.4/10 | 2026-06-29 03:32 |
| 904e70f9-b6b8-45… | Gate 3 | formula_repro |
How does intermediate-task training on English reasoning datasets affect zero-shot cross-lingual transfer performance on multilingual reason…
COUNTEREXAMPLE HUNTER: 7.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-06-29 03:32 |
| c7b84a89-c495-4c… | Gate 3 | formula_repro |
Does the effectiveness of English intermediate-task training for zero-shot cross-lingual transfer diminish when the target tasks require com…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 5.2/10 · REPLICATION ATTACKER: 6.5/10
|
6.3/10 | 2026-06-28 21:32 |
| 1f52dd3c-cfb8-46… | Gate 3 | formula_repro |
How does intermediate-task training on English NLI datasets affect zero-shot cross-lingual transfer performance on the XNLI subset of the XT…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.2/10
|
6.1/10 | 2026-06-28 21:32 |
| 489a4f4f-f6b1-48… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on XTREME diminish as the scale of the underlying pretrained multilingual …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-06-28 21:32 |
| 92b9f309-fda3-49… | Gate 3 | formula_repro |
To what extent does intermediate-task training on non-English source languages improve zero-shot cross-lingual performance on XTREME-M bench…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-06-28 21:32 |
| 0a215135-fa03-44… | Gate 3 | formula_repro |
How does the difficulty of the intermediate task (measured by English accuracy) correlate with zero-shot cross-lingual transfer accuracy on …
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
4.7/10 | 2026-06-28 21:32 |
| 5e997a72-8a5a-4e… | Gate 3 | formula_repro |
How does fine-tuning on typologically diverse intermediate tasks compared to English-only fine-tuning affect zero-shot accuracy on non-Indo-…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-06-28 15:32 |
| ce0acf70-cbfd-48… | Gate 3 | formula_repro |
How does the robustness of multilingual intermediate-task training compare to English intermediate-task training in zero-shot cross-lingual …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-06-28 15:32 |
| c5c7d64a-0415-4c… | Gate 3 | formula_repro |
To what extent does the choice of intermediate task (e.g., natural language inference, question answering) affect zero-shot cross-lingual tr…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-06-28 15:32 |
| 7cad01cb-6305-4b… | Gate 3 | formula_repro |
How does the selection of diverse English intermediate tasks (e.g., NLI, SQuAD) impact zero-shot cross-lingual transfer accuracy when scalin…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-06-28 15:32 |
| ef261fbc-7801-45… | Gate 3 | formula_repro |
How does increasing the volume of target-language data during intermediate-task training affect zero-shot transfer accuracy on XTREME compar…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.3/10
|
6.4/10 | 2026-06-28 15:31 |
| 3e4804df-9588-4e… | Gate 3 | formula_repro |
How does the hybrid batch training strategy impact zero-shot retrieval accuracy on low-resource languages in the BEIR benchmark compared to …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 2.0/10
|
5.7/10 | 2026-06-28 15:31 |
| c76d817d-0ca7-4f… | Gate 3 | formula_repro |
Which specific English intermediate language understanding tasks most effectively improve zero-shot cross-lingual robustness against adversa…
COUNTEREXAMPLE HUNTER: 3.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.1/10 | 2026-06-28 09:29 |
| 5079a0ff-1aa2-44… | Gate 3 | formula_repro |
What is the impact of varying batch composition ratios (monolingual vs. cross-lingual vs. multilingual) on zero-shot retrieval performance f…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 1.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.5/10 | 2026-06-28 09:29 |
| 8a86a6c1-062f-49… | Gate 3 | formula_repro |
How does the performance of self-supervised speech models pre-trained on Flemish Dutch compare to multilingual models when evaluated on cros…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.0/10 | 2026-06-28 09:29 |
| cf498b2a-0202-41… | Gate 3 | formula_repro |
To what extent does English intermediate-task training improve alignment consistency across non-English languages in zero-shot transfer sett…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.3/10
|
6.4/10 | 2026-06-28 03:29 |
| 140fb5e4-404c-48… | Gate 3 | formula_repro |
How does model size scaling impact the effectiveness of English intermediate-task transfer for zero-shot cross-lingual understanding on XTRE…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.3/10 | 2026-06-28 03:28 |
| b7c4a827-afa6-41… | Gate 3 | formula_repro |
Does the performance gap between self-supervised models pre-trained on Flemish Dutch and those pre-trained on English persist when evaluated…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.8/10 | 2026-06-27 21:26 |
| 1fb48f93-ff7c-42… | Gate 3 | formula_repro |
How does intra-inter knowledge distillation for mutual guidance between intent and slot tasks impact zero-shot cross-lingual accuracy on the…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 2.5/10
|
6.5/10 | 2026-06-27 21:26 |
| 56f9e4cf-821a-4c… | Gate 3 | formula_repro |
How does adapter-based fine-tuning compare to full model fine-tuning in cross-lingual NER accuracy on low-resource languages within the Wiki…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-06-27 21:26 |
| 6a8f6d7d-8f29-46… | Gate 3 | formula_repro |
To what extent does increasing out-of-domain unlabeled data volume degrade entity boundary precision in low-resource cross-lingual NER when …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 3.5/10
|
6.2/10 | 2026-06-27 21:26 |
| dcd682f4-a15a-42… | Gate 3 | formula_repro |
How does the robustness of Flemish Dutch self-supervised speech models compare to English pre-trained models when evaluated on noisy or adve…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.2/10
|
7.6/10 | 2026-06-27 15:23 |
| 32a1bd56-3c64-4e… | Gate 3 | formula_repro |
How does the cross-lingual transfer performance of self-supervised speech models pre-trained on Flemish Dutch compare to models pre-trained …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.8/10 | 2026-06-27 15:23 |
| 9f9bd4ab-6dc0-48… | Gate 3 | formula_repro |
How does extending the I²KD mutual guidance mechanism to vision-language models impact zero-shot cross-lingual intent detection accuracy on …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 2.5/10
|
5.8/10 | 2026-06-27 15:23 |
| f2da0cac-9cb0-42… | Gate 3 | formula_repro |
How does the robustness of teacher-student learning for cross-lingual NER compare to direct model transfer when evaluated on adversarial exa…
COUNTEREXAMPLE HUNTER: 8.0/10 · CITATION AUDITOR: 1.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-06-27 15:23 |
| fed29ffb-1707-42… | Gate 3 | formula_repro |
What is the impact of script variation on the robustness of cross-lingual transfer learning for Named Entity Recognition in low-resource Afr…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 3.5/10
|
6.8/10 | 2026-06-27 09:21 |
| e5e6823d-e830-49… | Gate 3 | formula_repro |
What is the impact of varying the ratio of monolingual to cross-lingual training samples in hybrid batch training on retrieval accuracy for …
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
5.7/10 | 2026-06-27 03:20 |
| 9bbb129b-1a6c-4e… | Gate 3 | formula_repro |
What is the comparative word error rate of self-supervised speech models pre-trained on low-resource Flemish Dutch versus fine-tuned English…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.0/10 | 2026-06-27 03:20 |
| 34f4e234-8df7-4e… | Gate 3 | formula_repro |
How does the zero-shot cross-lingual transfer performance of English intermediate-task trained models compare to models fine-tuned on multil…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-26 21:16 |
| 0d26825f-d8b7-4d… | Gate 3 | formula_repro |
Does the effectiveness of English intermediate-task training generalize to cross-domain multilingual tasks (e.g., biomedical, social media) …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.2/10
|
5.4/10 | 2026-06-26 21:16 |
| dbd33eed-4372-4d… | Gate 3 | formula_repro |
How does the performance of projection-based cross-lingual NER models compare to multilingual BERT models when evaluated on the Wikiner benc…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-26 21:12 |
| db4ac7b7-9dc8-4e… | Gate 3 | formula_repro |
How does multilingual intermediate-task fine-tuning compare to English-only fine-tuning in improving zero-shot cross-lingual transfer perfor…
COUNTEREXAMPLE HUNTER: 7.3/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.2/10
|
6.0/10 | 2026-06-26 21:12 |
| 01c3ec9e-55c7-4b… | Gate 3 | formula_repro |
What is the impact of varying the size and diversity of the English intermediate-task dataset on zero-shot cross-lingual transfer performanc…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.7/10 | 2026-06-26 21:12 |
| ce40ff48-e0ea-48… | Gate 3 | formula_repro |
How does intermediate-task training on multimodal vision-language datasets impact zero-shot cross-lingual accuracy on the XTREME benchmark c…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-06-26 21:12 |
| 793f877f-97da-46… | Gate 3 | formula_repro |
What is the effect of bilingual lexicon quality on the performance gains of dense retrievers trained with artificially code-switched data fo…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-06-26 15:10 |
| 9a7eb4c9-0e89-45… | Gate 3 | formula_repro |
How does fine-tuning with artificially code-switched data compare to multilingual pre-training in terms of cross-lingual generalization, as …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-06-26 15:10 |
| 9c28d35f-1996-47… | Gate 3 | formula_repro |
How does training on artificially code-switched data impact the zero-shot retrieval accuracy of multilingual dense retrievers on low-resourc…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.3/10 | 2026-06-26 15:10 |
| 40713f75-699f-4f… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual ranking models vary when trained on artificial code-switched data generated from bilingu…
COUNTEREXAMPLE HUNTER: 3.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.2/10
|
6.7/10 | 2026-06-26 15:10 |
| f0bee894-2752-48… | Gate 3 | formula_repro |
How does the robustness of zero-shot cross-lingual retrieval models trained on code-switched data compare to models trained on monolingual d…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-06-26 09:06 |
| 3bd9a949-71f4-46… | Gate 3 | formula_repro |
To what extent can zero-shot cross-lingual retrieval performance be improved by combining artificially code-switched data with multilingual …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-06-26 09:06 |
| adf01cf8-07b2-4a… | Gate 3 | formula_repro |
How does the use of artificially code-switched datasets impact the alignment of multilingual embeddings in models like LASER or LaBSE when e…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.3/10 · REPLICATION ATTACKER: 8.5/10
|
6.4/10 | 2026-06-26 09:06 |
| a25cea1a-2cde-46… | Gate 3 | formula_repro |
What is the impact of varying the proportion of artificial code-switching in the training data on the zero-shot cross-lingual retrieval accu…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.5/10 | 2026-06-26 09:05 |
| 4009340c-b38f-4d… | Gate 3 | formula_repro |
How does the zero-shot cross-lingual retrieval performance of models pre-trained on synthetic code-switching data compare to those trained o…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-06-26 09:05 |
| 08fbbf18-3433-4a… | Gate 3 | formula_repro |
To what extent does English intermediate-task training enhance the robustness of zero-shot cross-lingual transfer on XTREME-R tasks under do…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-06-26 09:05 |
| 7d0326ed-9d79-47… | Gate 3 | formula_repro |
How does intermediate fine-tuning on domain-specific corpora in non-English languages compare to multilingual intermediate tasks for zero-sh…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
8.7/10 | 2026-06-26 03:05 |
| 9376f177-c147-4b… | Gate 3 | formula_repro |
How does intermediate-task training on multilingual tasks (e.g., XNLI) compare to English-only intermediate tasks for improving zero-shot cr…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
6.2/10 | 2026-06-26 03:05 |
| f2b3fecb-4faf-46… | Gate 3 | formula_repro |
What is the impact of English intermediate-task training on the robustness of multilingual models to typological distance in zero-shot cross…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-06-26 03:05 |
| a8f8ac1f-a2e5-4d… | Gate 3 | formula_repro |
How does intermediate-task training on English reasoning datasets impact the adversarial robustness of mBERT in zero-shot cross-lingual sent…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 3.5/10
|
6.2/10 | 2026-06-26 03:05 |
| 1f89687c-4808-41… | Gate 3 | formula_repro |
Does combining English intermediate-task training with data augmentation techniques further improve zero-shot cross-lingual transfer perform…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-06-26 03:05 |
| 565ae775-5db3-43… | Gate 3 | formula_repro |
Does the benefit of English intermediate-task training for zero-shot cross-lingual transfer on XTREME-R persist when scaling to larger multi…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 6.5/10
|
4.3/10 | 2026-06-26 03:05 |
| a80809ed-afe5-4d… | Gate 3 | formula_repro |
How does the choice of intermediate-task complexity (e.g., simple vs. complex NLI tasks) affect zero-shot cross-lingual transfer accuracy on…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.1/10 | 2026-06-26 03:05 |
| 9589102c-5a18-40… | Gate 3 | formula_repro |
How does the performance of multilingual intermediate-task training compare to English intermediate-task training in zero-shot cross-lingual…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
8.7/10 | 2026-06-25 21:04 |
| b62dc7bf-b8a5-4a… | Gate 3 | formula_repro |
What is the impact of task diversity in intermediate fine-tuning on the robustness of zero-shot cross-lingual transfer across morphologicall…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.5/10 | 2026-06-25 21:04 |
| c9a9c158-f744-47… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on multilingual models scale with model size when evaluated on low-resourc…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.4/10 | 2026-06-25 21:04 |
| cbf93ce4-9990-4f… | Gate 3 | formula_repro |
How does the choice of intermediate task difficulty (e.g., easy vs. hard tasks) affect the robustness of zero-shot cross-lingual transfer pe…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 6.2/10 · REPLICATION ATTACKER: 7.2/10
|
6.9/10 | 2026-06-25 21:04 |
| 68c4910b-a25b-48… | Gate 3 | formula_repro |
How does interleaved multilingual intermediate-task training affect zero-shot transfer accuracy on XTREME compared to sequential English-onl…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-06-25 21:04 |
| 3b227090-6de8-4b… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on non-English XTREME targets scale with increasing pretrained model size …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-06-25 21:04 |
| f44cdeae-87c7-40… | Gate 3 | formula_repro |
How does scaling the size of the intermediate-task training dataset influence the robustness of zero-shot cross-lingual transfer on XTREME b…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-06-25 21:04 |
| f943782f-d8f9-47… | Gate 3 | formula_repro |
How does domain-specific intermediate-task training (e.g., medical or legal English tasks) impact zero-shot cross-lingual transfer robustnes…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 9.5/10
|
7.3/10 | 2026-06-25 15:04 |
| 42d802c3-92eb-46… | Gate 3 | formula_repro |
Can contrastive learning objectives in multimodal models improve alignment across languages and enhance robustness in zero-shot cross-lingua…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-06-25 15:04 |
| a3d46fa0-7214-4c… | Gate 3 | formula_repro |
Does increasing the diversity of intermediate task domains further improve zero-shot cross-lingual transfer performance on XTREME compared t…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-06-25 15:03 |
| f13442e4-7260-4d… | Gate 3 | formula_repro |
Does the effectiveness of English intermediate-task training on zero-shot cross-lingual transfer vary by the semantic similarity between the…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-06-25 15:02 |
| baa6b180-aa0e-48… | Gate 3 | formula_repro |
What is the impact of multilingual intermediate-task training (e.g., using XNLI or PAWS-X) on zero-shot cross-lingual transfer performance c…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-25 15:01 |
| 4f2dd2dc-f707-4d… | Gate 2 | unknown | To what extent does cross-lingual transfer via sequential fine-tuning improve F1-scores for euphemism detection in low-resource languages li… | - | 2026-06-25 09:49 |
| 8796b9ed-bf52-40… | Gate 3 | formula_repro |
How does the hybrid batch training strategy scale to multilingual language models beyond Bloom-7B (e.g., Bloom-176B) in terms of performance…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-25 02:57 |
| d97f943c-8f43-45… | Gate 3 | formula_repro |
How does varying word alignment noise levels in annotation projection impact the F1-score of cross-lingual NER compared to adversarial train…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 2.0/10
|
6.8/10 | 2026-06-25 02:57 |
| b74eccf8-e2e8-44… | Gate 3 | formula_repro |
How does the combination of adversarial alignment and projection-based data transfer compare to multilingual language models in terms of F1-…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.0/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-25 02:57 |
| dc422de2-bc8d-45… | Gate 3 | formula_repro |
To what extent does incorporating synthetic noise into unlabeled target data improve the alignment accuracy of cross-lingual NER student mod…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-25 02:57 |
| 04e4c078-c4e1-4c… | Gate 3 | formula_repro |
How does incorporating image-text pairs into annotation projection impact cross-lingual NER accuracy on XTREME-R low-resource languages comp…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.5/10
|
6.2/10 | 2026-06-25 02:56 |
| 74c5e2d3-579a-40… | Gate 3 | formula_repro |
How do different projection-based cross-lingual NER techniques compare in performance (F1 scores) when applied to languages with varying deg…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-25 02:56 |
| bebc437b-ec91-4d… | Gate 3 | formula_repro |
To what extent does increasing the volume of out-of-domain unlabeled target data improve the robustness of cross-lingual NER alignment in lo…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-25 02:56 |
| 8f099e7d-5931-49… | Gate 3 | formula_repro |
How does the robustness of mT5 and XLM-R in zero-shot cross-lingual retrieval compare to monolingual models trained on code-switched data wh…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-24 20:55 |
| a2b30927-4495-4c… | Gate 3 | formula_repro |
How does training on artificially code-switched data affect the zero-shot retrieval accuracy of multilingual dense retrievers on low-resourc…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 8.5/10
|
5.0/10 | 2026-06-24 20:55 |
| b14419d3-45dc-49… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare to models fine-tuned…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.0/10
|
8.8/10 | 2026-06-24 20:54 |
| c18e8cb4-8ed1-4d… | Gate 3 | formula_repro |
To what extent does the choice of pre-trained multilingual language model (e.g., mBERT, XLM-R, Bloom) affect the zero-shot cross-lingual tra…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.0/10
|
8.9/10 | 2026-06-24 20:53 |
| d63d5df0-8dac-42… | Gate 3 | formula_repro |
How does increasing the diversity of source languages compared to scaling unlabeled target corpus size affect few-shot cross-lingual NER acc…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 9.5/10
|
6.7/10 | 2026-06-24 20:53 |
| 1d4f99a8-293e-44… | Gate 3 | formula_repro |
How does the zero-shot cross-lingual retrieval accuracy of XLM-R and mBART compare when trained on code-switched data from multiple low-reso…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.0/10 | 2026-06-24 14:49 |
| 11bc427f-036c-47… | Gate 3 | formula_repro |
Can the use of dynamic word embeddings in zero-shot cross-lingual retrieval models improve robustness to domain shifts when trained on code-…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.7/10 | 2026-06-24 14:49 |
| 13d5b5a6-03c0-43… | Gate 3 | formula_repro |
How does training on artificially code-switched data affect the robustness of cross-lingual retrieval models against adversarial perturbatio…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.8/10 | 2026-06-24 14:47 |
| 3731e0ca-e6c8-41… | Gate 3 | formula_repro |
Can artificially code-switched training data mitigate the performance drop in zero-shot retrieval when queries and documents are in differen…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-06-24 14:46 |
| 01fb6973-7294-49… | Gate 3 | formula_repro |
How does training on artificially code-switched data generated via contextual embedding replacement compare to bilingual lexicon substitutio…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.5/10
|
6.8/10 | 2026-06-24 14:45 |
| 8f65f54a-11e7-4b… | Gate 3 | formula_repro |
How does the ratio of artificial code-switched tokens in training data affect the zero-shot performance of multilingual models like XLM-R on…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-06-24 14:44 |
| 5a7f0e70-2cdb-43… | Gate 3 | formula_repro |
What is the effect of scaling the proportion of non-English WebFAQ training samples on the retrieval performance of dense models for low-res…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 8.5/10
|
9.0/10 | 2026-06-24 14:43 |
| 323ff867-bf47-40… | Gate 3 | formula_repro |
How do zero-shot cross-lingual retrieval models trained on adversarial code-switched data perform on low-resource languages compared to high…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 8.5/10
|
8.9/10 | 2026-06-24 08:40 |
| 846293ff-5857-41… | Gate 3 | formula_repro |
How does training on artificially code-switched data affect zero-shot cross-lingual retrieval accuracy on scientific domain benchmarks like …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 3.5/10
|
6.3/10 | 2026-06-24 08:39 |
| 13d968ff-4b6f-40… | Gate 3 | formula_repro |
How does the robustness of artificially code-switched training data compare to multilingual pre-training when evaluated on zero-shot cross-l…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.3/10
|
7.6/10 | 2026-06-24 08:37 |
| ca574310-2657-4d… | Gate 3 | formula_repro |
Does the effectiveness of artificially code-switched training for zero-shot cross-lingual retrieval generalize to multimodal retrieval tasks…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.0/10 | 2026-06-24 08:36 |
| 33ff433a-693d-4a… | Gate 3 | formula_repro |
How does increasing model scale from 3B to 70B parameters impact zero-shot cross-lingual retrieval accuracy on XNLI when trained on artifici…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.7/10 | 2026-06-24 08:35 |
| 482c6454-2d82-47… | Gate 3 | formula_repro |
What is the effect of increasing the proportion of code-switched tokens in artificially generated training data on the zero-shot cross-lingu…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 7.5/10
|
8.3/10 | 2026-06-24 08:35 |
| 752f06b2-d5e2-44… | Gate 3 | formula_repro |
How does the choice of different bilingual lexicon sources (e.g., Wiktionary, machine-translated lexicons) impact the XNLI accuracy of zero-…
COUNTEREXAMPLE HUNTER: 8.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.7/10 | 2026-06-24 08:32 |
| 7d07e894-2c18-4c… | Gate 3 | formula_repro |
How does the choice of intermediate task semantic similarity affect the zero-shot cross-lingual transfer performance of multilingual models …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-06-24 02:30 |
| 3c3f51ae-90de-46… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual transfer in XTREME tasks vary when using intermediate-task training on high-resource non…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-06-24 02:29 |
| e5bed67d-9b36-40… | Gate 3 | formula_repro |
Does scaling the number of English intermediate tasks improve zero-shot cross-lingual transfer performance on the XTREME benchmark, and if s…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
8.7/10 | 2026-06-23 20:17 |
| 4e115898-45d1-46… | Gate 3 | formula_repro |
What is the robustness of zero-shot cross-lingual transfer performance on XTREME when intermediate-task training is applied to multimodal mo…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.4/10 | 2026-06-23 20:17 |
| 3ed1f46b-0547-46… | Gate 3 | formula_repro |
How does the order of fine-tuning on multiple English intermediate tasks affect zero-shot cross-lingual generalization accuracy on XTREME-R …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-23 20:17 |
| 21afb709-9ce2-4a… | Gate 3 | formula_repro |
Does the semantic similarity between English intermediate tasks and target XTREME-R tasks correlate with the magnitude of zero-shot cross-li…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-23 20:17 |
| 78a5ee74-3e2e-47… | Gate 3 | formula_repro |
How does simultaneous multi-task intermediate training compare to sequential training in improving zero-shot cross-lingual accuracy on XTREM…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-23 20:17 |
| 269f9df2-a33a-4a… | Gate 3 | formula_repro |
Does multi-task intermediate training on English QA degrade robustness to adversarial perturbations in zero-shot cross-lingual transfer comp…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 6.5/10
|
7.7/10 | 2026-06-23 20:17 |
| 16c701c9-a709-4f… | Gate 3 | formula_repro |
How does the hybrid batch training strategy scale with increasing model size (e.g., 6B to 175B parameters) and how does it affect the trade-…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.0/10
|
8.7/10 | 2026-06-23 20:17 |
| 4f8a7f27-6337-41… | Gate 3 | formula_repro |
Does simultaneous optimization of monolingual and cross-lingual objectives degrade reasoning capabilities in multilingual language models wh…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.0/10
|
5.8/10 | 2026-06-23 20:16 |
| dba1f4be-411c-43… | Gate 2 | unknown | What is the effect of varying the number of languages in sequential fine-tuning (e.g., 2 vs. 5 languages) on zero-shot cross-lingual accurac… | - | 2026-06-23 19:59 |
| ea483193-13d8-4b… | Gate 3 | formula_repro |
Does scaling the size of the intermediate-task dataset in low-resource languages improve zero-shot cross-lingual transfer robustness on XTRE…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-23 14:15 |
| 5ae21486-68ce-46… | Gate 3 | formula_repro |
To what extent does the typological distance between the English intermediate task language and the target language impact zero-shot transfe…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 6.5/10
|
8.2/10 | 2026-06-23 14:12 |
| 0ee93aea-a2e7-4f… | Gate 3 | formula_repro |
How does intermediate-task training on low-resource language families compare to English-only intermediate training for zero-shot transfer p…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.4/10 | 2026-06-23 14:11 |
| 5d9340a1-a95c-49… | Gate 3 | formula_repro |
Does the benefit of English intermediate-task transfer persist in multilingual models pre-trained on non-English corpora when evaluated on l…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-23 14:10 |
| 9a41c15b-ab00-44… | Gate 3 | formula_repro |
To what extent does intermediate-task training on English data improve robustness against adversarial perturbations in zero-shot cross-lingu…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.0/10 | 2026-06-23 14:09 |
| 47f4b735-26be-43… | Gate 3 | formula_repro |
Do multilingual intermediate tasks improve zero-shot cross-lingual transfer performance more consistently than English-only tasks across dif…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.1/10 | 2026-06-23 14:07 |
| 3a5f1363-74bc-48… | Gate 3 | formula_repro |
Does the effectiveness of English intermediate-task training for cross-lingual transfer in XTREME generalize to other low-resource language …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-06-23 14:05 |
| 5bdc0bd8-492f-47… | Gate 3 | formula_repro |
How does the choice of intermediate task (e.g., NLI, QQP, STS) affect zero-shot cross-lingual transfer performance on XTREME-R when combined…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-06-23 14:02 |
| a2761cf3-95af-4c… | Gate 3 | formula_repro |
How does intermediate-task diversity impact zero-shot cross-lingual transfer accuracy on morphologically complex languages within the XTREME…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-23 13:59 |
| 51155009-141e-4c… | Gate 2 | unknown | How does scaling model parameters from 100M to 10B affect the marginal gain of English intermediate-task training on zero-shot XTREME-R clas… | - | 2026-06-23 10:31 |
| 775aaa06-1f8c-47… | Gate 2 | unknown | What is the comparative effect of English intermediate-task fine-tuning versus direct target-language fine-tuning on alignment metrics for u… | - | 2026-06-23 08:00 |
| 2df779eb-945d-47… | Gate 3 | formula_repro |
How does the proposed hybrid batch training ratio affect the zero-shot cross-lingual retrieval performance of CLIP and ALBEF models on the X…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 2.0/10
|
6.8/10 | 2026-06-23 07:53 |
| e175141f-5207-41… | Gate 3 | formula_repro |
How does the simultaneous optimization of monolingual, cross-lingual, and multilingual retrieval performance scale with model size, as measu…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-23 07:53 |
| ee6a5ccb-c10c-44… | Gate 3 | formula_repro |
How does the hybrid batch training strategy proposed in this paper compare to adaptive batch sampling methods in terms of zero-shot cross-li…
COUNTEREXAMPLE HUNTER: 2.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.0/10
|
7.0/10 | 2026-06-23 07:52 |
| cad75bcb-3a65-4e… | Gate 3 | formula_repro |
To what extent does the performance of XLM-R models on zero-shot retrieval tasks degrade when trained exclusively on cross-lingual data vs. …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-23 07:51 |
| 1f8b202e-6e3c-44… | Gate 3 | formula_repro |
How does varying the proportion of hard negatives in hybrid monolingual and cross-lingual training batches affect zero-shot retrieval accura…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.0/10
|
6.2/10 | 2026-06-23 07:51 |
| fff97eca-c74f-47… | Gate 3 | formula_repro |
How does the hybrid batch training strategy compare to contrastive learning methods in terms of zero-shot cross-lingual retrieval performanc…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.5/10
|
6.0/10 | 2026-06-23 07:50 |
| bcf0c72c-6f91-48… | Gate 3 | formula_repro |
What is the comparative performance of the hybrid batch training strategy versus separate monolingual and cross-lingual fine-tuning on the B…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-23 07:49 |
| 0d8af95b-a8cd-41… | Gate 3 | formula_repro |
How does intermediate-task training on syntactic parsing versus semantic similarity affect zero-shot cross-lingual accuracy on the XNLI logi…
COUNTEREXAMPLE HUNTER: 2.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 6.5/10
|
6.1/10 | 2026-06-23 01:48 |
| cc1b6aca-6c16-4e… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on XTREME diminish when applied to multilingual models pretrained with non…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-23 01:48 |
| 02b8352e-256b-46… | Gate 3 | formula_repro |
Does intermediate-task training on diverse source languages enhance robustness to domain shift in zero-shot cross-lingual transfer compared …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-23 01:47 |
| 43d6d435-79a5-49… | Gate 3 | formula_repro |
How does the duration of English intermediate-task training affect the consistency of zero-shot cross-lingual transfer performance across ty…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.9/10 · REPLICATION ATTACKER: 9.2/10
|
9.0/10 | 2026-06-23 01:46 |
| 95c9d425-7129-46… | Gate 3 | formula_repro |
Does intermediate-task training on typologically distant source languages improve zero-shot accuracy on XTREME classification tasks more tha…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-23 01:46 |
| 28c5c8cc-fe5c-43… | Gate 3 | formula_repro |
What is the impact of intermediate-task training on the cross-lingual robustness of LLMs, measured by performance variance across XTREME-R l…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
5.6/10 | 2026-06-23 01:46 |
| a96596a8-30ab-44… | Gate 3 | formula_repro |
How does intermediate-task training on typologically diverse low-resource languages compare to English intermediate tasks for zero-shot tran…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
8.3/10 | 2026-06-23 01:46 |
| d3d4270d-7638-4f… | Gate 3 | formula_repro |
How does increasing pretraining model size affect the calibration error gap between English and morphologically rich languages in zero-shot …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-06-22 19:45 |
| ede19014-53f9-4d… | Gate 3 | formula_repro |
Does scaling the size of the English intermediate task dataset yield diminishing returns for zero-shot cross-lingual transfer compared to us…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-06-22 19:45 |
| 89290bc8-1e89-43… | Gate 3 | formula_repro |
How does intermediate-task training on multilingual NLI datasets compare to English-only NLI tasks for zero-shot transfer performance on the…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.5/10 | 2026-06-22 19:45 |
| 213a64f5-4833-48… | Gate 3 | formula_repro |
Does the benefit of English intermediate-task training for cross-lingual transfer on XTREME diminish when the target language belongs to a d…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 2.5/10
|
6.1/10 | 2026-06-22 19:45 |
| c409f8b0-7c55-46… | Gate 3 | formula_repro |
How does increasing the size of English intermediate-task training data affect zero-shot accuracy on non-English subsets of the XTREME bench…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.2/10
|
7.4/10 | 2026-06-22 19:45 |
| 81449533-4e2b-4c… | Gate 3 | formula_repro |
How does the zero-shot cross-lingual transfer performance of models with English intermediate-task training compare to multilingual intermed…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-22 19:45 |
| ef90b2e5-8bba-40… | Gate 3 | formula_repro |
Does intermediate-task training on English NLU tasks improve zero-shot cross-lingual transfer performance on XGLUE when applied to mT5-XXL c…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-22 19:44 |
| 9993c53e-950a-41… | Gate 3 | formula_repro |
What is the impact of varying the ratio of code-switched to monolingual data in training on the zero-shot performance of multilingual dense …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-22 13:37 |
| 32bad463-9e51-48… | Gate 3 | formula_repro |
To what extent does training on artificially code-switched data with varying switch frequencies improve zero-shot retrieval accuracy for low…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-22 13:35 |
| 6524d153-b03d-47… | Gate 3 | formula_repro |
How does domain-specific data augmentation influence the robustness of the hybrid batch training strategy across monolingual, cross-lingual,…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.0/10
|
8.8/10 | 2026-06-22 13:35 |
| d8e97f9c-f574-47… | Gate 3 | formula_repro |
What is the impact of combining contrastive loss with the hybrid batch training strategy on multilingual semantic alignment, as measured by …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-22 13:34 |
| 56c3a0d7-8522-40… | Gate 3 | formula_repro |
How does the hybrid batch training strategy perform when scaled to Bloom-7B compared to mT5-base on the XNLI benchmark, and what is the resu…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.0/10
|
9.1/10 | 2026-06-22 13:34 |
| ef9ef58e-ade8-41… | Gate 3 | formula_repro |
To what extent does the simultaneous optimization of monolingual and cross-lingual objectives in the hybrid batch strategy improve the R-pre…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 8.5/10
|
9.0/10 | 2026-06-22 13:33 |
| 30694f6d-6574-4a… | Gate 3 | formula_repro |
Does optimizing the hybrid batch composition for simultaneous monolingual and cross-lingual objectives degrade zero-shot reasoning performan…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.0/10
|
9.2/10 | 2026-06-22 13:31 |
| 997c4da8-1a01-42… | Gate 3 | formula_repro |
To what extent does the proposed hybrid batch training strategy improve cross-lingual retrieval robustness on the XTREME benchmark compared …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-06-22 13:30 |
| 93321a8a-451d-4c… | Gate 2 | unknown | How does the scaling of model size (e.g., 3B to 175B parameters) influence the effectiveness of zero-shot cross-lingual retrieval when train… | - | 2026-06-22 08:31 |
| d97e1003-3991-40… | Gate 3 | formula_repro |
What is the effect of varying the proportion of code-switched tokens (e.g., 10%, 30%, 50%) in artificially generated training data on the re…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-06-22 07:25 |
| c5c00c72-ac77-4c… | Gate 3 | formula_repro |
How does the choice of bilingual lexicon source (e.g., WordNet, LIDA) impact the robustness of zero-shot cross-lingual retrieval models trai…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 3.5/10
|
6.5/10 | 2026-06-22 07:25 |
| a8ae9a79-9828-43… | Gate 3 | formula_repro |
What is the impact of scaling artificially code-switched training data size on zero-shot cross-lingual retrieval performance, measured by ac…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.5/10 | 2026-06-22 07:24 |
| 0ec149b5-f9cd-43… | Gate 3 | formula_repro |
Can fine-tuning zero-shot cross-lingual retrieval models on artificially code-switched data generated by multilingual embeddings improve rob…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 6.5/10
|
4.7/10 | 2026-06-22 07:24 |
| 3038c26d-8ede-4e… | Gate 3 | formula_repro |
How does the robustness of zero-shot cross-lingual retrieval models trained on code-switched data compare to those trained on monolingual da…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-06-22 07:24 |
| 8fe795f4-6614-44… | Gate 3 | formula_repro |
What is the impact of varying the proportion of code-switched tokens in artificial training data on the zero-shot cross-lingual retrieval ac…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-06-22 07:24 |
| 547ecb20-fa89-43… | Gate 3 | formula_repro |
How does the performance of cross-lingual dense passage retrievers compare when trained on artificially code-switched data versus multilingu…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-06-22 07:24 |
| 1c4d3bc7-4ddd-4d… | Gate 3 | formula_repro |
How does the generalization capability of cross-lingual retrieval models trained on artificially code-switched data perform on out-of-domain…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.5/10
|
7.0/10 | 2026-06-22 07:24 |
| b6469e92-c832-49… | Gate 3 | formula_repro |
What is the impact of varying the size of the English intermediate-task dataset on the robustness of cross-lingual transfer to low-resource …
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 9.5/10
|
7.5/10 | 2026-06-22 01:24 |
| 19ac4394-8f06-47… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training in zero-shot cross-lingual transfer diminish when scaling to larger pretra…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 4.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-06-22 01:24 |
| d28d8022-0827-42… | Gate 3 | formula_repro |
Does English intermediate-task training maintain robustness in zero-shot cross-lingual transfer across diverse language families in XTREME w…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.3/10 | 2026-06-22 01:24 |
| 0b14d79d-59df-4d… | Gate 3 | formula_repro |
What is the correlation between the semantic similarity of English intermediate tasks and target non-English tasks in XTREME and the resulti…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-06-22 01:24 |
| fa66da7b-d27d-42… | Gate 3 | formula_repro |
How does incorporating cross-lingual query generation into dense retrieval models affect performance on zero-shot cross-lingual transfer tas…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 3.5/10
|
6.0/10 | 2026-06-21 19:23 |
| e1d57bf6-8fdd-42… | Gate 3 | formula_repro |
How does the zero-shot cross-lingual voice cloning performance of ASR-guided flow-matching TTS models compare to diffusion-based architectur…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.0/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-06-21 19:23 |
| 11c231b9-3e2b-40… | Gate 3 | formula_repro |
To what extent does the proposed hybrid training strategy improve cross-lingual retrieval robustness for rare scripts in XM3600 compared to …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.7/10 | 2026-06-21 19:23 |
| e0a7faed-c699-4d… | Gate 2 | unknown | How does the performance of XLM-R and mBERT compare in simultaneous multilingual fine-tuning for euphemism detection across high-resource an… | - | 2026-06-21 17:51 |
| f3323655-945b-44… | Gate 3 | formula_repro |
Does the WEAM pre-training strategy improve robustness against typological divergence in zero-shot cross-lingual transfer performance on XNL…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-21 13:23 |
| 3c02a222-c341-4b… | Gate 3 | formula_repro |
How does the hybrid batch training strategy's performance on zero-shot cross-lingual retrieval tasks compare to state-of-the-art models like…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.0/10
|
9.2/10 | 2026-06-21 13:23 |
| d7bbfc63-d748-49… | Gate 3 | formula_repro |
How does scaling the size of bilingual lexicons used to generate code-switched data affect the cross-lingual retrieval accuracy on XNLI when…
COUNTEREXAMPLE HUNTER: 8.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.0/10 | 2026-06-21 13:22 |
| 4c195f37-7769-4f… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare to multilingual mode…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.8/10 | 2026-06-21 13:22 |
| 7400f84d-f2c0-46… | Gate 3 | formula_repro |
How does the performance of cross-lingual retrieval models trained on artificially code-switched data compare to those trained on synthetic …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-21 13:22 |
| 59fee39e-f0e9-4e… | Gate 3 | formula_repro |
Does the performance gain from code-switched training on high-resource language pairs generalize effectively to low-resource language settin…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.3/10 | 2026-06-21 13:21 |
| c57d9aff-89f5-4b… | Gate 3 | formula_repro |
Does scaling the ratio of code-switched tokens in training data improve zero-shot retrieval accuracy on the MIRACL benchmark, and what is th…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-06-21 13:21 |
| 53cecda3-587c-4c… | Gate 3 | formula_repro |
How does the robustness of neural ranking models trained on synthetic code-switched data compare to models trained on original monolingual d…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 6.5/10
|
8.0/10 | 2026-06-21 07:21 |
| 24fd756c-a4d5-40… | Gate 3 | formula_repro |
Does scaling the volume of artificially code-switched training data yield diminishing returns in cross-lingual retrieval performance on the …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-21 07:21 |
| 3b3d07f6-2a83-45… | Gate 3 | formula_repro |
How does training on artificially code-switched data compare to multilingual contrastive learning in improving zero-shot cross-lingual retri…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-21 07:20 |
| 30d06151-f9d9-40… | Gate 3 | formula_repro |
Does incorporating artificially code-switched queries during training improve the alignment of embedding spaces for low-resource languages i…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.7/10 | 2026-06-21 07:20 |
| 8d61dab6-340e-40… | Gate 3 | formula_repro |
How does training on artificially code-switched data affect the zero-shot retrieval accuracy of multilingual LLMs on typologically dissimila…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.2/10
|
6.9/10 | 2026-06-21 07:20 |
| ed299d26-54ba-4e… | Gate 3 | formula_repro |
Does the benefit of English intermediate-task training for zero-shot cross-lingual transfer persist when evaluated on reasoning-heavy subset…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-21 07:20 |
| 53dd88c9-20f5-42… | Gate 3 | formula_repro |
How does the performance of multilingual intermediate-task training compare to English-only intermediate-task training when evaluated on XTR…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 3.5/10
|
7.2/10 | 2026-06-21 07:20 |
| f03bca7f-0a50-41… | Gate 3 | formula_repro |
To what extent does the typological distance between English and target low-resource languages modulate the performance improvements gained …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.7/10 | 2026-06-21 07:20 |
| 760fbc69-65b1-4f… | Gate 3 | formula_repro |
How does the scaling of model size affect the performance gain from English intermediate-task training in zero-shot cross-lingual transfer, …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.6/10 | 2026-06-21 01:19 |
| 9a6f19a5-f2e6-4c… | Gate 3 | formula_repro |
To what extent does English intermediate-task training improve cross-lingual reasoning capabilities on multilingual benchmarks compared to d…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.2/10
|
6.9/10 | 2026-06-21 01:19 |
| 5571973f-50bb-48… | Gate 3 | formula_repro |
Does intermediate-task training on domain-specific multilingual datasets improve robustness to domain shift in zero-shot cross-lingual trans…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.2/10 | 2026-06-21 01:19 |
| a80c506e-d2f8-4a… | Gate 3 | formula_repro |
What is the impact of scaling the number of intermediate language-understanding tasks on zero-shot cross-lingual transfer performance for lo…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-21 01:18 |
| 2290ba9d-c1c0-46… | Gate 3 | formula_repro |
What is the impact of intermediate-task training on low-resource languages in the XTREME benchmark when using models pretrained on both Engl…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.5/10
|
6.2/10 | 2026-06-21 01:18 |
| 2af78e0e-7fba-4d… | Gate 3 | formula_repro |
How does the effectiveness of English intermediate-task training for zero-shot cross-lingual transfer compare to multilingual intermediate-t…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-21 01:18 |
| 8544463f-7e19-46… | Gate 3 | formula_repro |
How does the performance of multilingual intermediate-task training on low-resource languages compare to English intermediate tasks when eva…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
4.9/10 | 2026-06-21 01:18 |
| bd43d6f2-a6c8-44… | Gate 3 | formula_repro |
How does the performance of intermediate-task training sequences compare to continuous pretraining on a multilingual corpus in zero-shot cro…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 2.5/10
|
6.1/10 | 2026-06-21 01:18 |
| 345ee2d0-d119-47… | Gate 3 | formula_repro |
Does multi-task intermediate training on diverse English NLU tasks improve robustness against typological divergence more effectively than s…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 6.5/10
|
8.0/10 | 2026-06-20 19:17 |
| 13542541-aac3-44… | Gate 3 | formula_repro |
What is the impact of English intermediate-task training on the alignment stability of multilingual encoders when evaluated on adversarial p…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-20 19:17 |
| d7a818b5-ede1-48… | Gate 3 | formula_repro |
Does the performance gain from English intermediate-task training on XTREME scale with increasing pretraining model size across diverse low-…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-06-20 19:17 |
| 7f4a6491-9864-42… | Gate 3 | formula_repro |
How does English intermediate-task training affect zero-shot cross-lingual robustness on XTREME tasks with synthetic code-switching noise co…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-20 19:17 |
| aa533e6e-e763-4c… | Gate 3 | formula_repro |
How does intermediate-task training on English reasoning datasets affect zero-shot cross-lingual performance on the XCOPA and XNLI subsets o…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.2/10
|
6.1/10 | 2026-06-20 19:17 |
| 48a8684b-3452-40… | Gate 3 | formula_repro |
Does the order of intermediate-task fine-tuning (sequential vs. concurrent) influence the robustness of multilingual alignment in zero-shot …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.1/10 | 2026-06-20 19:16 |
| a76c6aed-2013-42… | Gate 3 | formula_repro |
How does the choice of English intermediate-task difficulty (e.g., low vs. high complexity) affect zero-shot cross-lingual transfer performa…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 9.2/10
|
6.7/10 | 2026-06-20 19:16 |
| 62dad032-b551-46… | Gate 3 | formula_repro |
How does the choice of intermediate task complexity (e.g., easy vs. hard language understanding tasks) affect zero-shot cross-lingual transf…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 9.0/10
|
7.1/10 | 2026-06-20 19:15 |
| 4ffd234e-3144-46… | Gate 3 | formula_repro |
Does multilingual intermediate-task training on XTREME-R outperform monolingual English training in few-shot cross-lingual transfer across l…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
6.1/10 | 2026-06-20 19:15 |
| 5f20fd9d-3000-49… | Gate 2 | unknown | Does the integration of synthetic code-switched data improve the robustness of zero-shot cross-lingual retrieval models against adversarial … | - | 2026-06-20 16:59 |
| c29d0b1d-1c55-45… | Gate 2 | unknown | How does hybrid batch training for monolingual and cross-lingual objectives impact zero-shot retrieval accuracy on the BEIR benchmark compar… | - | 2026-06-20 16:57 |
| a10e30bc-cc3b-42… | Gate 3 | formula_repro |
How does training on artificially code-switched data affect the robustness of zero-shot cross-lingual retrieval models across low-resource l…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 3.5/10
|
6.8/10 | 2026-06-20 13:15 |
| 88b977da-978a-4b… | Gate 3 | formula_repro |
How does the granularity of bilingual lexicons (e.g., word-level vs. phrase-level) impact the effectiveness of artificially code-switched tr…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.8/10 | 2026-06-20 13:15 |
| 9293eb3f-c631-43… | Gate 3 | formula_repro |
What is the impact of artificially code-switched training data on the robustness of cross-lingual retrieval models evaluated on the PAWS-X d…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-06-20 13:15 |
| 108e3723-cd80-4b… | Gate 3 | formula_repro |
How does the quality of bilingual lexicons impact the performance of zero-shot cross-lingual retrieval models on the BEIR benchmark when eva…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.4/10 | 2026-06-20 13:14 |
| 4cb997d4-121c-4d… | Gate 3 | formula_repro |
Does training on artificially code-switched data improve zero-shot cross-lingual retrieval recall on the MIRACL benchmark compared to monoli…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-06-20 13:14 |
| 420aa4ee-b924-40… | Gate 3 | formula_repro |
What is the effect of increasing the amount of artificially code-switched training data on the robustness of zero-shot cross-lingual retriev…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-06-20 13:14 |
| bb51b78d-f1e5-48… | Gate 3 | formula_repro |
How does the transfer learning performance of self-supervised speech models pre-trained on Flemish Dutch compare to other low-resource langu…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
5.3/10 | 2026-06-20 07:14 |
| 1249d225-3960-40… | Gate 3 | formula_repro |
How does the noise level in automatically induced bilingual lexicons affect the nDCG@10 and MAP scores of zero-shot cross-lingual retrievers…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-06-19 19:12 |
| 2e7a84f4-8bf1-4d… | Gate 3 | formula_repro |
To what extent does English intermediate-task training enhance zero-shot reasoning capabilities on multilingual benchmarks like XTREME-R for…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-06-19 19:12 |
| d41c927e-b869-40… | Gate 3 | formula_repro |
How does intermediate-task training on non-English source languages compare to English-only intermediate training for zero-shot cross-lingua…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-19 19:12 |
| 44a960a9-7f95-41… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare to models fine-tuned…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.7/10 | 2026-06-19 13:09 |
| 05adb989-b8ea-44… | Gate 3 | formula_repro |
Does training on artificially code-switched datasets improve the robustness of zero-shot cross-lingual retrievers against query-document lan…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
5.7/10 | 2026-06-19 13:06 |
| da82e7da-f16a-41… | Gate 3 | formula_repro |
Does the hybrid batch strategy improve zero-shot cross-lingual retrieval robustness on the XTD benchmark compared to standard multilingual c…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-19 13:06 |
| 2bb9a5bd-57ad-44… | Gate 3 | formula_repro |
How does training on artificially code-switched data affect zero-shot cross-lingual performance on the XNLI benchmark compared to standard m…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.8/10 | 2026-06-19 13:06 |
| 3b872189-de14-45… | Gate 3 | formula_repro |
How does training on artificially code-switched data affect the zero-shot retrieval accuracy of multilingual dense retrievers on the Lasers …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.5/10 | 2026-06-19 13:06 |
| 5ff9bb46-9d08-44… | Gate 3 | formula_repro |
To what extent does training on artificially code-switched data improve zero-shot cross-lingual retrieval robustness on XTREME-R when querie…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.7/10 | 2026-06-19 13:05 |
| 6ef56955-1505-47… | Gate 3 | formula_repro |
How does the proportion of code-switched tokens in synthetic training data correlate with the accuracy drop of zero-shot cross-lingual ranke…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.0/10
|
8.8/10 | 2026-06-19 13:05 |
| 73b8a1d9-aee2-48… | Gate 3 | formula_repro |
How does training on artificially code-switched data affect the robustness of zero-shot cross-lingual rankers against adversarial noise comp…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.5/10 | 2026-06-19 13:03 |
| 07f2d90e-a6ad-40… | Gate 2 | unknown | Does training on artificially code-switched data improve zero-shot cross-lingual retrieval performance for low-resource languages not includ… | - | 2026-06-19 12:46 |
| 7507c92e-3093-47… | Gate 3 | formula_repro |
Does integrating CausalMixFT during fine-tuning improve the robustness of tabular foundation models against adversarial perturbations in low…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.0/10 | 2026-06-19 07:00 |
| da92f3e0-4d5b-44… | Gate 3 | formula_repro |
How do dense RGB-D SLAM systems utilizing 3D Gaussian representations compare to neural implicit methods in terms of memory consumption and …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-06-19 07:00 |
| 3d9090b2-81c7-4d… | Gate 3 | formula_repro |
How do vision-language models perform in cross-domain robustness evaluations when tested on perturbed multimodal benchmarks from domains lik…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.5/10 | 2026-06-19 07:00 |
| 0ff89691-6471-49… | Gate 3 | formula_repro |
How does the trade-off between model size and latency compare between OpenPangu-7B-MLA and smaller prosody-exclusive models when deployed on…
COUNTEREXAMPLE HUNTER: 8.0/10 · CITATION AUDITOR: 4.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.0/10 | 2026-06-19 07:00 |
| 86059cb7-4f35-4c… | Gate 3 | formula_repro |
What is the impact of cross-lingual transfer from English pre-trained speech models versus monolingual Flemish pre-training on phoneme recog…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-06-19 07:00 |
| 0cb45f8d-b6ae-45… | Gate 3 | formula_repro |
How does the addition of self-supervised pre-training objectives in zero-shot cross-lingual SLU models affect slot-filling accuracy on the M…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 2.5/10
|
6.5/10 | 2026-06-19 07:00 |
| 36c4669f-6437-43… | Gate 3 | formula_repro |
What is the effect of varying the size of the monolingual training set on the intent detection performance of zero-shot cross-lingual SLU mo…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 3.0/10
|
6.3/10 | 2026-06-19 07:00 |
| d245e830-fc51-49… | Gate 3 | formula_repro |
What is the impact of varying the code-switching ratio in training data on the retrieval performance degradation when query and document lan…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-06-19 01:00 |
| a5ab3b78-a7cb-48… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare to models fine-tuned…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.7/10 | 2026-06-19 01:00 |
| 5e61f043-f854-46… | Gate 3 | formula_repro |
What is the impact of varying the ratio of code-switched tokens in artificially generated training data on the robustness (measured by accur…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.8/10 | 2026-06-19 00:59 |
| a984d373-612a-42… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual rankers trained on artificially code-switched data compare to models fine-tuned on multi…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-06-19 00:59 |
| 963c284c-e2a8-44… | Gate 3 | formula_repro |
How does increasing the proportion of code-switched tokens in the training data affect the robustness of zero-shot cross-lingual retrieval m…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 8.5/10
|
5.0/10 | 2026-06-19 00:59 |
| 643ef90b-e1b6-4f… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare to models fine-tuned…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 7.5/10
|
6.0/10 | 2026-06-19 00:59 |
| 05f4c33e-153e-4c… | Gate 3 | formula_repro |
Does scaling the multilingual pre-trained model size improve precision@k in zero-shot cross-lingual retrieval when using the proposed hybrid…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.0/10
|
9.2/10 | 2026-06-18 18:59 |
| f988e6e1-5aab-4f… | Gate 3 | formula_repro |
Can scaling the hybrid batch training method to larger multilingual models (e.g., XLM-R or mT5) further enhance zero-shot cross-lingual retr…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-18 18:59 |
| 6a1d78e6-6b7d-44… | Gate 3 | formula_repro |
How does domain-adaptive fine-tuning of Flemish Dutch self-supervised speech models impact word error rate on CommonVoice compared to cross-…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-18 18:56 |
| 52c35878-037b-40… | Gate 3 | formula_repro |
What is the comparative effect of multi-task intermediate training versus single large-task training on reasoning capabilities within zero-s…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-18 18:54 |
| 9b3a269a-5094-47… | Gate 3 | formula_repro |
Does combining diverse intermediate tasks improve robustness in zero-shot cross-lingual transfer on XTREME-R more effectively than training …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.0/10 | 2026-06-18 18:53 |
| 880c758a-54df-4a… | Gate 3 | formula_repro |
How does hybrid batch training impact zero-shot cross-lingual retrieval accuracy on XNLI compared to monolingual fine-tuning across varying …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.5/10
|
7.0/10 | 2026-06-18 18:53 |
| faa8fed9-29ac-4e… | Gate 3 | formula_repro |
Does the synergistic hybrid batch training approach improve cross-lingual retrieval robustness on the MIRACL benchmark under domain shift co…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-18 18:51 |
| fd7ab5fe-aab9-48… | Gate 3 | formula_repro |
What is the impact of hybrid batch training on the scaling behavior of zero-shot retrieval performance across varying model sizes within the…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 2.0/10
|
6.5/10 | 2026-06-18 18:51 |
| 2e00ae69-fcc4-4e… | Gate 3 | formula_repro |
How does hybrid batch training for simultaneous monolingual and cross-lingual retrieval impact zero-shot performance on low-resource languag…
COUNTEREXAMPLE HUNTER: 8.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 4.5/10
|
7.3/10 | 2026-06-18 18:50 |
| 25c348c6-4569-45… | Gate 3 | formula_repro |
Does the synergistic optimization of monolingual and cross-lingual objectives in hybrid batch training improve retrieval performance on long…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-18 18:48 |
| a9afc122-af0a-4c… | Gate 3 | formula_repro |
What is the impact of varying the proportion of code-switched tokens in artificially generated training data on the robustness of zero-shot …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.7/10 | 2026-06-18 12:48 |
| 96a5be5a-20a9-4d… | Gate 3 | formula_repro |
Does training on artificially code-switched data improve the robustness of retrieval models against language mismatch errors in queries and …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-06-18 12:48 |
| d4730df9-5abd-4e… | Gate 3 | formula_repro |
What is the impact of varying bilingual lexicon coverage on the zero-shot cross-lingual retrieval performance of code-switched trained model…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-18 12:48 |
| 82718235-0464-46… | Gate 3 | formula_repro |
How does the cross-lingual retrieval accuracy of models trained on artificially code-switched data compare to full multilingual pretraining …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.2/10
|
9.2/10 | 2026-06-18 12:48 |
| 304f71e8-cfc0-47… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models improve when trained on artificially code-switched data generated from …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-06-18 12:47 |
| 5b801b53-702c-47… | Gate 3 | formula_repro |
How does the hybrid batch training strategy impact the zero-shot cross-lingual retrieval accuracy of larger multimodal models (e.g., PaLI, B…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-18 12:47 |
| 681c606e-0739-4d… | Gate 3 | formula_repro |
How does the scaling of model size (e.g., small, base, large) interact with the hybrid batch training strategy in terms of zero-shot cross-l…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-18 12:47 |
| fed390d7-0f30-4f… | Gate 3 | formula_repro |
Can the hybrid batch training strategy be adapted to improve zero-shot cross-lingual retrieval performance in low-resource language settings…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-18 12:46 |
| 69ee4e5c-067c-4b… | Gate 3 | formula_repro |
How does the scaling of intermediate-task dataset size affect the degradation of zero-shot cross-lingual transfer performance on the XTREME …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-18 06:45 |
| a9ba0423-52f5-40… | Gate 3 | formula_repro |
What is the impact of intermediate-task training on the robustness of zero-shot cross-lingual transfer to low-resource languages within the …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-18 06:45 |
| dc84b96b-753f-46… | Gate 3 | formula_repro |
Does the choice of multilingual intermediate tasks (e.g., language-agnostic vs. language-specific) impact the robustness of zero-shot cross-…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-18 06:45 |
| bff5aa00-46f0-42… | Gate 3 | formula_repro |
How does the performance of multilingual intermediate-task training compare to English intermediate-task training on the XTREME-R benchmark,…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.2/10 | 2026-06-18 06:44 |
| b6aec771-9353-47… | Gate 3 | formula_repro |
How does the hybrid batch training strategy impact zero-shot retrieval accuracy on low-resource MIRACL language pairs compared to dedicated …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 6.5/10
|
5.7/10 | 2026-06-18 06:44 |
| 82a8b49a-ca92-48… | Gate 3 | formula_repro |
How does fine-tuning Flemish Dutch self-supervised speech models with domain adaptation techniques affect word error rate on the CommonVoice…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.5/10 | 2026-06-18 06:44 |
| 16fd9c48-0866-43… | Gate 3 | formula_repro |
What is the impact of structural causal model fidelity on the downstream classification accuracy of fine-tuned tabular foundation models in …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-18 06:43 |
| 27b34f6d-ce74-49… | Gate 3 | formula_repro |
What is the effect of domain-specific vs. general-domain code-switched data on zero-shot cross-lingual retrieval performance in multilingual…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-06-18 00:41 |
| 706377b3-65ea-44… | Gate 3 | formula_repro |
How does the robustness of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare across different lang…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-06-18 00:41 |
| 6724432f-b734-46… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare to those trained on …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.2/10 | 2026-06-18 00:41 |
| 5eeb7591-2ce2-45… | Gate 3 | formula_repro |
How does the retrieval accuracy per training token of models trained on artificially code-switched data compare to full multilingual pretrai…
COUNTEREXAMPLE HUNTER: 6.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.5/10 | 2026-06-18 00:41 |
| 572586b4-d231-44… | Gate 3 | formula_repro |
How does the lexical coverage ratio of bilingual dictionaries used for artificial code-switching correlate with zero-shot cross-lingual retr…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.2/10 · REPLICATION ATTACKER: 6.5/10
|
7.1/10 | 2026-06-18 00:41 |
| b539628a-b9a7-4a… | Gate 3 | formula_repro |
How does hybrid batch training impact zero-shot retrieval recall@10 on the MIRACL benchmark for low-resource languages compared to monolingu…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-17 18:40 |
| 8d5a627d-354d-4d… | Gate 3 | formula_repro |
How does the hybrid batch training strategy impact zero-shot retrieval accuracy on unseen low-resource language pairs when evaluated on the …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-17 18:39 |
| 16872a52-c7b2-4c… | Gate 3 | formula_repro |
How does fine-tuning on naturally occurring code-switched corpora (e.g., LINCS or NLPCC) compare to fine-tuning on artificially code-switche…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-06-17 18:39 |
| a068b0e0-7ac3-4d… | Gate 3 | formula_repro |
Does the synergistic hybrid batch training strategy improve zero-shot cross-lingual retrieval accuracy for languages with varying typologica…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.0/10
|
8.8/10 | 2026-06-17 18:38 |
| e864d5a7-01cf-48… | Gate 3 | formula_repro |
To what extent do self-supervised speech models pre-trained on Flemish Dutch generalize to low-resource dialects compared to English pre-tra…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-06-17 18:38 |
| 4d287f11-5746-4c… | Gate 3 | formula_repro |
Does fine-tuning English pre-trained speech models on limited Flemish data yield comparable robustness to noise as models pre-trained exclus…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.5/10 | 2026-06-17 18:36 |
| 52cce802-988b-44… | Gate 3 | formula_repro |
How does scaling the model size of TSDiff impact its performance on cross-domain time series forecasting benchmarks (e.g., UCR archive) comp…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-17 18:36 |
| 69e07f68-b58d-4d… | Gate 3 | formula_repro |
What is the impact of simultaneous monolingual and cross-lingual objective optimization on the generalization capability of multilingual enc…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-17 12:34 |
| 8aad928b-e796-47… | Gate 3 | formula_repro |
How does hybrid batch training affect zero-shot cross-lingual retrieval accuracy on low-resource language pairs in the MIRACL benchmark comp…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-17 12:34 |
| 96ee3e4c-90e2-47… | Gate 3 | formula_repro |
How does varying the ratio of monolingual to cross-lingual training examples in hybrid batches affect the performance trade-off between NQ a…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.0/10
|
8.8/10 | 2026-06-17 12:34 |
| c1c990b6-70f0-44… | Gate 3 | formula_repro |
What is the impact of simultaneous monolingual, cross-lingual, and multilingual optimization on the retrieval performance of transformer mod…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
6.3/10 | 2026-06-17 12:31 |
| efef605f-428e-4d… | Gate 3 | formula_repro |
How does the synergistic hybrid batch training strategy compare to standard multilingual fine-tuning in terms of zero-shot cross-lingual ret…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-17 12:31 |
| 5e54cc94-1461-48… | Gate 3 | formula_repro |
Can integrating domain-specific monolingual data (e.g., legal, medical) into hybrid batch training improve zero-shot retrieval accuracy on X…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-06-17 12:28 |
| 45ea17e5-1565-49… | Gate 3 | formula_repro |
Can the model-agnostic nature of SafeCoDe be validated across different multimodal architectures (e.g., LLaVA, Qwen-VL) by comparing their s…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.7/10 | 2026-06-17 12:28 |
| d4118170-4257-4f… | Gate 3 | formula_repro |
How does the hybrid batch training strategy compare to language-specific adapter modules in improving zero-shot cross-lingual retrieval accu…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.0/10
|
9.1/10 | 2026-06-17 06:27 |
| f36a2df0-3772-49… | Gate 3 | formula_repro |
What is the impact of varying the degree of artificial code-switching in training data on the robustness of zero-shot cross-lingual retrieva…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-17 06:27 |
| 8acc7c8a-b815-43… | Gate 3 | formula_repro |
Does the hybrid batch training strategy improve zero-shot cross-lingual retrieval performance on non-English language pairs in XM3600 compar…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.5/10 | 2026-06-17 06:26 |
| 479ce4ad-ca7a-40… | Gate 3 | formula_repro |
Can intermediate-task training on English reasoning datasets mitigate cross-lingual performance degradation in low-resource languages in the…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-17 06:26 |
| cf208549-5033-41… | Gate 3 | formula_repro |
Does multilingual intermediate-task training improve zero-shot transfer accuracy on XTREME-R domain-specific subsets compared to English-onl…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 6.5/10
|
8.2/10 | 2026-06-17 06:26 |
| 5ad5a3a9-7cbb-4f… | Gate 3 | formula_repro |
What is the impact of varying the size and linguistic diversity of the English intermediate-task corpus on the degradation of zero-shot tran…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.0/10
|
8.4/10 | 2026-06-17 06:26 |
| c64ed49c-d8b1-42… | Gate 3 | formula_repro |
How does cross-lingual query generation augmentation impact the adversarial robustness of dense retrieval models against paraphrase attacks …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.5/10
|
8.1/10 | 2026-06-17 06:26 |
| 9cbec7df-52ae-46… | Gate 3 | formula_repro |
Does pretraining zero-shot cross-lingual retrieval models on artificially code-switched data improve robustness to language divergence in qu…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-17 00:25 |
| 791b2ca6-fdad-44… | Gate 3 | formula_repro |
Does the hybrid batch training strategy improve zero-shot cross-lingual retrieval robustness in low-resource language settings for multimoda…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-17 00:25 |
| c3b1fbef-bb8a-4c… | Gate 3 | formula_repro |
Does the hybrid batch training strategy improve retrieval performance on the XOR benchmark compared to models optimized solely for cross-lin…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-17 00:25 |
| 1003c363-131a-43… | Gate 3 | formula_repro |
Does training on artificially code-switched data improve zero-shot cross-lingual retrieval performance on the MLQA benchmark compared to sta…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.2/10
|
6.4/10 | 2026-06-17 00:25 |
| 0195f05e-2ecb-43… | Gate 3 | formula_repro |
Does the hybrid batch training strategy proposed for information retrieval improve multimodal alignment accuracy on zero-shot cross-lingual …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.7/10 | 2026-06-17 00:25 |
| 54b44e18-53d0-4a… | Gate 3 | formula_repro |
Does intermediate-task training on English reasoning datasets improve zero-shot cross-lingual performance on the XCOPA and XNLI subsets of X…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-06-17 00:25 |
| 1fba2b0b-14c2-4f… | Gate 3 | formula_repro |
Can multilingual intermediate-task training outperform English-only intermediate training for zero-shot transfer on domain-specific subsets …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-17 00:24 |
| f9e78b9c-2598-43… | Gate 3 | formula_repro |
How does the size of the English intermediate-task corpus affect the degradation of zero-shot transfer accuracy on low-resource languages wi…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-17 00:24 |
| f9cfb858-1284-49… | Gate 3 | formula_repro |
How does the performance of cross-lingual dense retrieval systems using query-augmented passage representations compare to those using multi…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.0/10
|
9.1/10 | 2026-06-17 00:24 |
| 8416a433-960a-43… | Gate 3 | formula_repro |
What is the impact of scaling the size of the synthetic dataset generated by CausalMixFT on the fine-tuning performance of tabular foundatio…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.8/10 | 2026-06-17 00:24 |
| 36871c14-b8df-4a… | Gate 3 | formula_repro |
How does cross-lingual query generation augmentation affect the adversarial robustness of dense retrieval models against paraphrase attacks …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-06-16 18:21 |
| 006e8e4a-a201-4b… | Gate 3 | formula_repro |
How does the performance gap between high-resource and low-resource languages in cross-lingual retrieval models change when using different …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.2/10 | 2026-06-16 18:21 |
| b6d14fcd-ba14-42… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare to models fine-tuned…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.2/10
|
5.7/10 | 2026-06-16 18:21 |
| ef1bc3a9-c7fa-44… | Gate 3 | formula_repro |
What is the impact of varying the proportion of code-switched terms in training data on the robustness of zero-shot cross-lingual retrieval …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-06-16 18:21 |
| b39d0c42-f3a9-4c… | Gate 3 | formula_repro |
How does the performance of cross-lingual query generation compare to multilingual contrastive learning (e.g., XLM-R, LasER) on the BEIR ben…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.7/10 | 2026-06-16 18:21 |
| 33238242-5dc0-45… | Gate 2 | unknown | How does the cross-lingual transfer performance of mE5 compare to other multilingual models like XLM-R or mBERT when pre-trained on monoling… | - | 2026-06-16 17:19 |
| 2556595d-7efa-45… | Gate 3 | formula_repro |
How does the performance of multilingual dense retrieval models compare on WebFAQ when trained with synthetic data augmentation versus human…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.5/10 | 2026-06-16 12:21 |
| fafdde58-463a-48… | Gate 3 | formula_repro |
How does the performance of Targeted Lexical Injection (TLI) with early-layer LoRA fine-tuning compare to full-parameter fine-tuning on the …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 7.5/10
|
8.2/10 | 2026-06-16 12:19 |
| 497fe53f-999a-48… | Gate 3 | formula_repro |
How does the pass@1 degradation of CodeT5 compare to JaCoText on MBPP Pro when subjected to semantic-preserving docstring perturbations vers…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.8/10 | 2026-06-16 12:19 |
| d3826925-d702-4e… | Gate 3 | formula_repro |
Can TLI early-layer LoRA fine-tuning improve cross-domain alignment in Lugha-Llama for low-resource Bantu languages, as evaluated by mAP sco…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-06-16 12:19 |
| d09dcdad-46aa-44… | Gate 3 | formula_repro |
What is the effect of Targeted Lexical Injection on cross-lingual alignment quality for Lugha-Llama when evaluated on semantic textual simil…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 6.5/10
|
7.5/10 | 2026-06-16 12:19 |
| c90a313e-3bcf-45… | Gate 3 | formula_repro |
How does early-layer LoRA with Targeted Lexical Injection impact zero-shot cross-lingual transfer accuracy on the XNLI benchmark for low-res…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 4.5/10
|
7.2/10 | 2026-06-16 12:19 |
| 647a30e8-8ab0-47… | Gate 3 | formula_repro |
To what extent does the depth of early-layer LoRA fine-tuning in TLI affect cross-lingual lexical alignment, as measured by LAS scores acros…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-06-16 06:19 |
| a8f56d30-b343-4a… | Gate 3 | formula_repro |
How do context-aware conversational models and sequence labeling approaches differ in zero-shot cross-lingual transfer accuracy for hate spe…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-06-16 06:18 |
| a63a1ae2-af59-40… | Gate 3 | formula_repro |
How does the scalability of CausalMixFT compare to other data augmentation methods (e.g., SMOTE, GAN-based augmentation) when fine-tuning ta…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-16 06:18 |
| 75b0a51f-2504-44… | Gate 3 | formula_repro |
Can SCM-based synthetic augmentation reduce the validation data requirements for early stopping in fine-tuning, as measured by the stability…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-16 06:18 |
| 3dd55559-98c7-47… | Gate 3 | formula_repro |
Do parameter-efficient fine-tuning methods like LoRA maintain instance segmentation performance on COCO when applied to other transformer ba…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 0.0/10
|
5.5/10 | 2026-06-16 06:18 |
| b451d079-8c4b-4d… | Gate 3 | formula_repro |
How does the integration of CausalMixFT-generated synthetic data affect the fine-tuning convergence speed and validation accuracy of tabular…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-06-16 06:18 |
| 069ce81a-1531-45… | Gate 3 | formula_repro |
Does CausalMixFT outperform diffusion-based data augmentation (e.g., DiffAugment) in terms of robustness to covariate shift when fine-tuning…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.0/10
|
9.1/10 | 2026-06-16 06:17 |
| a97c638e-389e-47… | Gate 3 | formula_repro |
How does integrating causal structure into TabPFN's synthetic data generation affect its performance on downstream task accuracy across diff…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.7/10 | 2026-06-16 06:17 |
| d7b2e2bf-4375-41… | Gate 2 | unknown | How do TimeGAN and VAE-generated synthetic financial time series compare in terms of robustness when used to evaluate the temporal reasoning… | - | 2026-06-16 01:04 |
| e0caf16c-fb4c-48… | Gate 3 | formula_repro |
How does varying the depth of LoRA adapter injection in Lugha-Llama affect cross-lingual alignment accuracy on low-resource Swahili-English …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-06-16 00:15 |
| bc889895-c22b-44… | Gate 3 | formula_repro |
To what extent does the combination of SFT and DPO degrade the zero-shot reasoning capabilities of OPT-350M on the Big-Bench Hard suite rela…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-06-16 00:14 |
| fcc53574-9cfa-44… | Gate 3 | formula_repro |
How does the reasoning accuracy of multimodal large language models compare to diffusion-based trajectory policies in dynamic task planning …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-06-16 00:14 |
| 852ac52e-42d2-4e… | Gate 3 | formula_repro |
How does the hybrid batch training strategy impact zero-shot cross-lingual retrieval accuracy on low-resource languages within the XQuAD ben…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-16 00:14 |
| f8fc038d-df35-4b… | Gate 3 | formula_repro |
How does the scaling of synthetic data diversity in tabular foundation model pretraining affect accuracy degradation under distributional sh…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.2/10 · REPLICATION ATTACKER: 8.5/10
|
7.9/10 | 2026-06-16 00:14 |
| f1cb0512-e3b3-44… | Gate 3 | formula_repro |
To what extent does incorporating causal priors via CausalMixFT improve out-of-distribution (OOD) robustness in tabular foundation models, a…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.5/10 | 2026-06-16 00:13 |
| 4efb4ac7-2558-4e… | Gate 3 | formula_repro |
How does the cross-lingual query generation approach compare to cross-lingual passage generation in terms of enhancing the alignment capabil…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.0/10
|
8.7/10 | 2026-06-15 18:13 |
| e4b97bd1-f023-4b… | Gate 3 | formula_repro |
What is the correlation between training data volume in WebFAQ 2.0 and zero-shot cross-lingual retrieval performance gaps across the 75 supp…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.5/10
|
6.8/10 | 2026-06-15 18:11 |
| b68d34da-d3d1-43… | Gate 3 | formula_repro |
Can synergistic optimization of monolingual and cross-lingual objectives reduce performance degradation on the XTREME retrieval benchmark fo…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 1.5/10
|
6.3/10 | 2026-06-15 18:10 |
| de04d7fd-23ac-49… | Gate 3 | formula_repro |
Does the hybrid batch training strategy improve zero-shot cross-lingual retrieval performance on downstream datasets like MIRACL or XNLI whe…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 7.5/10
|
8.5/10 | 2026-06-15 18:08 |
| 82ef2513-1ed2-43… | Gate 3 | formula_repro |
Does early-layer LoRA adaptation for lexical alignment in Lugha-Llama maintain zero-shot translation accuracy on morphologically rich Bantu …
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.8/10 | 2026-06-15 18:08 |
| 79f75ace-7cdb-4f… | Gate 3 | formula_repro |
How does the alignment of synthetic financial data generated by GANs versus VAEs influence the downstream performance of multimodal models i…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-15 18:07 |
| be269da2-4b74-48… | Gate 3 | formula_repro |
How does the noise level in automatically extracted bilingual lexicons impact the zero-shot cross-lingual retrieval accuracy of code-switche…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.5/10 · REPLICATION ATTACKER: 6.5/10
|
5.2/10 | 2026-06-15 18:07 |
| 1ce91234-09d8-44… | Gate 3 | formula_repro |
How does early-layer LoRA fine-tuning for lexical alignment in Lugha-Llama compare to full-parameter fine-tuning on zero-shot cross-lingual …
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.5/10 | 2026-06-15 18:07 |
| d78d2f8b-f3d8-45… | Gate 3 | formula_repro |
How does early-layer LoRA fine-tuning for lexical alignment in Lugha-Llama compare to full-parameter fine-tuning on cross-lingual retrieval …
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-06-15 12:06 |
| fb1c1363-5e72-48… | Gate 3 | formula_repro |
Does early-layer LoRA fine-tuning improve cross-lingual lexical alignment more effectively than full-model fine-tuning for low-resource Afri…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 7.5/10
|
7.4/10 | 2026-06-15 12:05 |
| 9f2d0d5b-7064-49… | Gate 3 | formula_repro |
How does fine-tuning dense retrieval models on WebFAQ's 47 million non-English pairs impact zero-shot cross-lingual transfer accuracy on the…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 3.5/10
|
7.2/10 | 2026-06-15 12:05 |
| c64e666b-968b-44… | Gate 3 | formula_repro |
How does training on artificially code-switched data compare to translate-train methods in improving zero-shot cross-lingual retrieval accur…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-06-15 12:05 |
| befe8d82-5024-48… | Gate 3 | formula_repro |
How does the performance of zero-shot cross-lingual retrieval models trained on artificially code-switched data compare to multilingual pret…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.2/10 | 2026-06-15 12:05 |
| 05a11663-8824-41… | Gate 3 | formula_repro |
How does hybrid batch training for simultaneous monolingual and cross-lingual optimization impact zero-shot retrieval accuracy on out-of-dom…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.5/10
|
6.2/10 | 2026-06-15 12:05 |
| e10a1785-aeeb-42… | Gate 3 | formula_repro |
Does training on artificially code-switched data improve cross-lingual retrieval precision compared to monolingual training when evaluated o…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.1/10 | 2026-06-15 12:05 |
| 28d7ecd5-529b-47… | Gate 3 | formula_repro |
Does training on artificially code-switched data improve cross-lingual robustness on the XQuAD benchmark when evaluated against standard mul…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 7.5/10
|
8.3/10 | 2026-06-15 12:05 |
| 324f98c4-488c-47… | Gate 3 | formula_repro |
Does intermediate-task training on domain-specific English corpora improve zero-shot transfer performance on multilingual domain subsets of …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 6.5/10
|
8.2/10 | 2026-06-15 12:05 |
| 367b042b-53cd-40… | Gate 3 | formula_repro |
How does the alignment between MIDI symbolic input and audio output in Tacotron-based models compare to that of neural source-filter wavefor…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-06-15 06:04 |
| 853e72b4-662b-4e… | Gate 3 | formula_repro |
What is the impact of TLI early-layer LoRA fine-tuning on the robustness of Lugha-Llama against adversarial lexical perturbations in low-res…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.7/10 | 2026-06-15 06:04 |
| e655d427-02d9-48… | Gate 3 | formula_repro |
How does early-layer LoRA adaptation in Lugha-Llama impact zero-shot cross-lingual retrieval accuracy on noisy Swahili-English datasets comp…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.8/10 | 2026-06-15 06:04 |
| b435d9ed-a132-4b… | Gate 3 | formula_repro |
How does early-layer LoRA fine-tuning for lexical injection compare to middle-layer adaptation in improving cross-lingual alignment scores o…
COUNTEREXAMPLE HUNTER: 3.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 9.5/10
|
6.7/10 | 2026-06-15 06:04 |
| c5a6b419-3445-44… | Gate 3 | formula_repro |
How does the token prioritization strategy in Vcc affect perplexity scores on the PG-19 benchmark compared to sparse attention patterns like…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 8.5/10
|
8.2/10 | 2026-06-15 06:04 |
| 757b4933-05c3-4b… | Gate 3 | formula_repro |
How does cross-lingual query generation compare to direct cross-lingual data training in terms of improving passage representation alignment…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-15 06:03 |
| 7ee1158c-08c4-4d… | Gate 3 | formula_repro |
Does augmenting passage representations with generated queries reduce the latency-throughput trade-off in cross-lingual dense retrieval syst…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.7/10 | 2026-06-15 06:03 |
| f149150d-181d-4e… | Gate 3 | formula_repro |
How does the robustness of zero-shot cross-lingual voice cloning in flow-matching TTS models vary when evaluated on noisy or adversarial inp…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.5/10 | 2026-06-15 06:03 |
| 67bf0cbf-613c-4e… | Gate 3 | formula_repro |
How does the combined SFT+DPO alignment strategy impact the reasoning accuracy of OPT-350M on complex multilingual queries relative to stand…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
6.2/10 | 2026-06-15 06:02 |
| 780db65d-2d7d-42… | Gate 3 | formula_repro |
To what extent does increasing the scale of the base language model mitigate the degradation in helpfulness scores observed in OPT-350M afte…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
4.7/10 | 2026-06-15 06:02 |
| 9013c53e-30cb-47… | Gate 3 | formula_repro |
How does early-layer LoRA adaptation for lexical alignment in Lugha-Llama compare to full fine-tuning in zero-shot cross-lingual transfer ac…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-06-15 00:02 |
| ae131cd2-d64a-42… | Gate 3 | formula_repro |
Does the latent cross-lingual alignment achieved via Targeted Lexical Injection in Lugha-Llama generalize to zero-shot machine translation p…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-06-15 00:01 |
| 8d6c501f-4675-4a… | Gate 3 | formula_repro |
Do auxiliary objectives with factorized latent dynamics improve sample efficiency in small-scale Video-JEPA training relative to standard jo…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 6.5/10
|
4.6/10 | 2026-06-15 00:01 |
| 9ade25b1-b2f8-40… | Gate 3 | formula_repro |
What is the effect of factorized latent dynamics auxiliary objectives on the transfer learning performance of Video-JEPA when evaluated on d…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.2/10 · REPLICATION ATTACKER: 6.5/10
|
7.1/10 | 2026-06-15 00:01 |
| afe7316d-46f8-46… | Gate 3 | formula_repro |
How does CLIP-TD's zero-shot transfer accuracy on domain-shifted vision-language tasks compare to standard CLIP fine-tuning methods?
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.5/10 | 2026-06-14 18:01 |
| db269ac8-afd2-4c… | Gate 3 | formula_repro |
How does the scaling of self-supervised pretraining data size affect the performance of few-shot meta-learners on language model benchmarks …
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-06-14 18:01 |
| 72958e36-47e4-4e… | Gate 3 | formula_repro |
What is the impact of integrating motion-image diffusion priors on the robustness of vision-language-action models against adversarial pertu…
COUNTEREXAMPLE HUNTER: 7.2/10 · CITATION AUDITOR: 3.2/10 · REPLICATION ATTACKER: 8.5/10
|
6.3/10 | 2026-06-14 18:01 |
| 96f30ae4-f497-45… | Gate 3 | formula_repro |
How does the cross-lingual voice cloning performance of flow-matching TTS models compare to diffusion-based TTS models when evaluated on uns…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.2/10 | 2026-06-14 18:01 |
| cad169a5-a675-41… | Gate 3 | formula_repro |
How does the performance of Targeted Lexical Injection (TLI) compare to full fine-tuning and adapter-based methods on the XTREME-R benchmark…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-14 18:01 |
| cf8d5c87-fe0e-4d… | Gate 3 | formula_repro |
How does hybrid batch training affect the zero-shot cross-lingual retrieval accuracy of mBERT on low-resource language pairs compared to mon…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.0/10
|
9.2/10 | 2026-06-14 17:59 |
| 463bd35f-feb7-45… | Gate 3 | formula_repro |
Does synergistic optimization of monolingual and cross-lingual objectives improve generalization to unseen language pairs in the BEIR zero-s…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 2.5/10
|
7.0/10 | 2026-06-14 17:58 |
| 964ed2b1-5e26-41… | Gate 3 | formula_repro |
How does hybrid batch training affect zero-shot retrieval accuracy on low-resource languages in the XTREME benchmark compared to dedicated m…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.0/10
|
9.2/10 | 2026-06-14 17:58 |
| a6b4186a-a702-48… | Gate 3 | formula_repro |
What is the comparative robustness of CausalMixFT-generated synthetic data against other data augmentation methods (e.g., GAN-based or diffu…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.7/10 | 2026-06-14 11:57 |
| 42f5bcad-eb35-4b… | Gate 3 | formula_repro |
To what extent does CausalMixFT fine-tuning improve the generalization accuracy of tabular foundation models under data scarcity compared to…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-06-14 11:57 |
| d9c74c0f-cb53-41… | Gate 3 | formula_repro |
How does the F1-score of multilingual transformer models compare to monolingual models when evaluated on code-mixed hate speech datasets wit…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.5/10 | 2026-06-14 11:57 |
| 5ec1afab-3eea-4c… | Gate 3 | formula_repro |
How does early-layer LoRA lexical injection compare to middle-layer adaptation in improving zero-shot cross-lingual retrieval accuracy for S…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-14 11:56 |
| 4e42da41-1273-43… | Gate 3 | formula_repro |
To what extent does Targeted Lexical Injection improve cross-lingual alignment scores on the XCOPA dataset for underrepresented Bantu langua…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.5/10
|
6.8/10 | 2026-06-14 11:56 |
| 5073185f-b8b7-47… | Gate 3 | formula_repro |
What is the impact of context window size on the retrieval-augmented generation performance of quantized LoRA-adapted models when evaluating…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.5/10 | 2026-06-14 11:56 |
| e486f159-c9b9-4f… | Gate 3 | formula_repro |
How does the alignment of multimodal embeddings (e.g., text and audio) in MUST-RAG affect the consistency and robustness of generated answer…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 9.5/10
|
6.8/10 | 2026-06-14 11:55 |
| 1fafcd2c-2be3-43… | Gate 3 | formula_repro |
How does the fidelity of structural causal models used for data augmentation impact the few-shot classification accuracy of fine-tuned tabul…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 3.0/10
|
6.5/10 | 2026-06-14 11:55 |
| e404ac75-ed39-43… | Gate 3 | formula_repro |
Can targeted lexical injection in Lugha-Llama achieve comparable zero-shot cross-lingual performance to MMPLMs like WMT21fb on clinical doma…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.0/10 | 2026-06-14 11:55 |
| 7b8d797a-9b95-44… | Gate 3 | formula_repro |
How does the use of causal data augmentation techniques like CausalMixFT compare to traditional data augmentation methods (e.g., SMOTE, GAN-…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 8.5/10
|
8.9/10 | 2026-06-14 05:55 |
| cf1474bf-5527-47… | Gate 3 | formula_repro |
What is the accuracy degradation of generalized zero-shot learning models under norm-bounded perturbations across unseen classes?
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-14 05:54 |
| a7211fa0-6c1a-43… | Gate 3 | formula_repro |
How does fine-tuning dense retrieval models on native multilingual WebFAQ data impact zero-shot cross-lingual retrieval accuracy on XQuAD co…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.0/10 | 2026-06-14 05:54 |
| 99730361-463c-4d… | Gate 3 | formula_repro |
Does early-layer LoRA fine-tuning improve zero-shot cross-lingual natural language inference accuracy for low-resource African languages com…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-06-14 05:54 |
| fdb047e5-d021-45… | Gate 3 | formula_repro |
How does the performance of dense retrieval models trained on WebFAQ compare to those trained on Wikipedia-based datasets like Natural Quest…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-14 05:53 |
| 3be16ffd-6b7d-4f… | Gate 3 | formula_repro |
How does the ratio of synthetic to real pretraining data impact the few-shot classification accuracy of multimodal video-language models on …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-14 05:53 |
| 038c8225-74bd-49… | Gate 3 | formula_repro |
How does contrastive pretraining objective selection impact cross-lingual retrieval accuracy for low-resource language pairs in the XTREME b…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.5/10
|
7.0/10 | 2026-06-14 05:53 |
| d1a71185-38af-4c… | Gate 3 | formula_repro |
Does hybrid batch training improve cross-domain generalization for multilingual retrieval models on unseen topics in low-resource languages …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-14 05:53 |
| 70f05cea-fe5d-40… | Gate 3 | formula_repro |
What is the impact of hybrid batch training on the scaling behavior of zero-shot retrieval accuracy when extending from low-resource to high…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-06-14 05:53 |
| 86a0376b-0fad-42… | Gate 3 | formula_repro |
What is the impact of mixed-precision inference (e.g., FP16 vs. BF16) on the efficiency-accuracy trade-off for long-context models like Long…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-06-14 05:53 |
| d7973fbb-05f8-47… | Gate 3 | formula_repro |
How does the zero-shot cross-lingual retrieval accuracy of a multilingual encoder pre-trained on WebFAQ's 47M non-English QA pairs compare t…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.0/10 · REPLICATION ATTACKER: 9.5/10
|
6.8/10 | 2026-06-13 23:53 |
| 3378e994-9dc5-4c… | Gate 3 | formula_repro |
Does scaling the proportion of non-English WebFAQ fine-tuning data improve retrieval latency and accuracy trade-offs for cross-lingual tasks…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-13 23:52 |
| 36a277bf-5c26-40… | Gate 3 | formula_repro |
What is the impact of incorporating visual modality into self-supervised learning for speech representations on the robustness of neural sou…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 4.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.7/10 | 2026-06-13 23:52 |
| 809b945c-726a-44… | Gate 3 | formula_repro |
What is the impact of mixed-dataset pretraining versus single-dataset pretraining on the robustness of Video-JEPA representations to tempora…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.8/10 | 2026-06-13 23:52 |
| da64baea-9f76-40… | Gate 3 | formula_repro |
What is the impact of varying the ratio of synthetic to real data in CausalMixFT on the fine-tuning performance of tabular foundation models…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-13 23:51 |
| b83881e6-6035-48… | Gate 3 | formula_repro |
What is the impact of causal data augmentation proportions on the sample efficiency and convergence speed of fine-tuning tabular foundation …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-13 23:51 |
| c90f70ac-13c1-4c… | Gate 3 | formula_repro |
What is the correlation between the fidelity of synthetic tabular samples generated via SCMs and the downstream fine-tuning performance of f…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 7.0/10
|
8.4/10 | 2026-06-13 23:51 |
| 0b01bfe2-6818-4c… | Gate 3 | formula_repro |
Does integrating causal structure into synthetic data generation improve the robustness of TabPFN against feature permutation compared to st…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.5/10 | 2026-06-13 23:51 |
| b744a045-e120-4e… | Gate 3 | formula_repro |
What is the impact of varying the proportion of causal synthetic data during fine-tuning on the robustness of tabular foundation models acro…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.7/10 | 2026-06-13 23:51 |
| 1ab47328-816a-4c… | Gate 3 | formula_repro |
How does the robustness of dense retrievers pretrained on WebFAQ compare to those trained on monolingual datasets when evaluated on adversar…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.8/10 | 2026-06-13 23:50 |
| e3b701fc-8dd1-45… | Gate 3 | formula_repro |
How does the performance of Video-JEPA models with factorized latent dynamics compare to non-factorized variants when evaluated on the Somet…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 8.5/10
|
5.0/10 | 2026-06-13 17:50 |
| 1613d088-356e-46… | Gate 3 | formula_repro |
What is the impact of mixed-dataset pretraining (UCF-101 + Something-Something V2 + ImageNet-100) on the accuracy of Video-JEPA models with …
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.2/10 | 2026-06-13 17:50 |
| 51b48b75-c13e-4d… | Gate 3 | formula_repro |
Does the robustness gained from Targeted Lexical Injection in Lugha-Llama generalize to code-switched social media text as measured by F1 sc…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-06-13 17:50 |
| fc099746-cbba-4e… | Gate 3 | formula_repro |
Does fine-tuning tabular foundation models with Structural Causal Model-based synthetic data improve generalization accuracy more than stand…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 8.5/10
|
6.5/10 | 2026-06-13 17:50 |
| 2003206f-03e3-4e… | Gate 3 | formula_repro |
Does combining ImageNet-100 with video datasets improve the domain robustness of self-supervised Video-JEPA representations on heterogeneous…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.3/10 | 2026-06-13 17:50 |
| a836297c-8d52-40… | Gate 3 | formula_repro |
What is the impact of varying the rank of LoRA matrices on cross-lingual alignment for Turkic languages when fine-tuned on early layers, eva…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.3/10 | 2026-06-13 17:50 |
| 469e9bad-7f2e-41… | Gate 3 | formula_repro |
How does the generalization performance of CausalMixFT compare to other data augmentation methods (e.g., Mixup, SMOTE) when fine-tuning tabu…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-06-13 17:50 |
| d74469d0-67aa-44… | Gate 3 | formula_repro |
How does fine-tuning dense retrieval models on WebFAQ's 47 million non-English pairs impact zero-shot cross-lingual transfer accuracy on the…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-13 17:50 |
| a76c6c0e-24bb-4e… | Gate 3 | formula_repro |
What is the comparative robustness of early-layer LoRA versus full-parameter fine-tuning for Lugha-Llama on cross-lingual natural language i…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-13 17:50 |
| 8df4041f-79f6-45… | Gate 3 | formula_repro |
How does the incorporation of auxiliary objectives in Video-JEPA models impact the robustness of learned representations when evaluated on o…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-06-13 11:49 |
| f33863b8-6dc3-46… | Gate 3 | formula_repro |
How does factorized latent dynamics in Video-JEPA compare to standard JEPA in cross-domain transfer accuracy from synthetic to real-world vi…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-13 11:49 |
| 063ca2b8-24d2-46… | Gate 3 | formula_repro |
What is the impact of varying the number of LoRA layers on cross-lingual lexical alignment in Lugha-Llama when benchmarked against the FLORE…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-13 11:49 |
| 9dadb434-2322-46… | Gate 3 | formula_repro |
How do bitwise neural networks with stochastic inference techniques perform in comparison to full-precision networks with Monte Carlo dropou…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.5/10 | 2026-06-13 11:49 |
| bebb97c4-3ac3-46… | Gate 3 | formula_repro |
What is the effect of the SFT+DPO alignment strategy on the helpfulness retention rate of OPT-350M when evaluated on the Anthropic Helpful-H…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-13 11:48 |
| 8a29c57f-aef4-47… | Gate 3 | formula_repro |
How does retrieval-augmented revision compare to adversarial training in improving Big-Vul detection accuracy for Llama-3.1-8B without requi…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.7/10 | 2026-06-13 11:48 |
| dbd73bd9-a579-40… | Gate 3 | formula_repro |
How does fine-tuning dense retrieval models on the non-English subset of WebFAQ impact cross-lingual zero-shot performance on TyDi QA compar…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 7.5/10
|
5.0/10 | 2026-06-13 11:48 |
| e890a0c6-75f2-41… | Gate 3 | formula_repro |
What is the impact of fine-tuning WebFAQ-pretrained dense retrieval models on downstream cross-lingual NLI tasks, as measured by XNLI accura…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-13 11:48 |
| 6910a60a-f0a2-40… | Gate 3 | formula_repro |
Do auxiliary factorized objectives in Video-JEPA improve few-shot learning performance on fine-grained video benchmarks relative to non-fact…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 9.5/10
|
8.1/10 | 2026-06-13 11:48 |
| 31ae88ad-ede0-43… | Gate 3 | formula_repro |
To what extent does Direct Preference Optimization enhance the robustness of counter-speech models against adversarial hate speech inputs co…
COUNTEREXAMPLE HUNTER: 4.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
7.7/10 | 2026-06-13 05:48 |
| d82913b1-e2c4-40… | Gate 3 | formula_repro |
How does retrieval diversity in music-specific RAG frameworks impact answer robustness against adversarial perturbations compared to general…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 8.5/10
|
9.0/10 | 2026-06-13 05:48 |
| 3cb37eff-ef87-4e… | Gate 3 | formula_repro |
How does the multimodal capture component in Expert Mind affect VQA accuracy on domain-specific datasets compared to text-only RAG baselines…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-13 05:47 |
| e447cada-12a9-4f… | Gate 3 | formula_repro |
What is the comparative effect of graph sparsity versus density on the F1-score performance of retrieval-augmented generation models in zero…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 5.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.0/10 | 2026-06-13 05:47 |
| fd93dc1c-d547-4d… | Gate 3 | formula_repro |
How does the MRR of cross-lingual dense retrieval models degrade on WebFAQ low-resource language families compared to high-resource ones whe…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.8/10 | 2026-06-13 05:47 |
| 840b7dfc-7587-47… | Gate 3 | formula_repro |
What is the impact of scaling the multilingual dense retriever model size (e.g., small vs. large) on retrieval performance across low-resour…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 9.5/10
|
6.8/10 | 2026-06-13 05:47 |
| 223a5aad-7c31-46… | Gate 3 | formula_repro |
To what extent does training on artificially code-switched data improve cross-lingual retrieval robustness for low-resource languages compar…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-13 05:47 |
| cf7bfdf5-1e02-40… | Gate 3 | formula_repro |
Does training dense retrievers on WebFAQ 2.0's bilingual aligned pairs improve zero-shot question answering accuracy on multilingual benchma…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-13 05:46 |
| 6811def3-b807-4f… | Gate 3 | formula_repro |
How does fine-tuning dense retrieval models on WebFAQ's non-English subsets impact zero-shot cross-lingual retrieval accuracy on the XTREME …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 7.5/10
|
8.7/10 | 2026-06-13 05:46 |
| 80d91e4a-f8da-4d… | Gate 3 | formula_repro |
What is the impact of injecting LoRA adapters exclusively into attention mechanisms versus feed-forward networks in Llama-3.2-3B on the late…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 9.5/10
|
7.5/10 | 2026-06-12 21:25 |
| 1b596d33-8278-4f… | Gate 3 | formula_repro |
What is the impact of fine-tuning CodeT5 with adversarial training on its semantic consistency and robustness accuracy in generalized zero-s…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-12 21:23 |
| 04a7cbf6-efb9-4b… | Gate 3 | formula_repro |
How does CausalMixFT compare to other data augmentation techniques (e.g., SMOTE, MixUp) in terms of fine-tuning robustness on tabular datase…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.5/10 | 2026-06-12 13:50 |
| 5e879d86-d825-41… | Gate 3 | formula_repro |
How does the ratio of synthetic-to-real data in CausalMixFT affect the F1 score variance of tabular foundation models on TabFact across mult…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-06-12 13:50 |
| d96ded43-64c0-44… | Gate 3 | formula_repro |
How does evidential deep learning with non-negative evidence constraints affect cross-modal retrieval accuracy on CLIP and ALBEF compared to…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.2/10 | 2026-06-12 13:50 |
| 6330a381-0e16-4e… | Gate 3 | formula_repro |
How does the data augmentation strategy used in scTab compare in effectiveness to other state-of-the-art data augmentation techniques when a…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 8.5/10
|
7.2/10 | 2026-06-12 07:43 |
| 63170d1a-bec2-45… | Gate 3 | formula_repro |
To what extent does the causal structure complexity (e.g., number of confounders or mediators) in the SCM used for CausalMixFT affect the ge…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 4.2/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-06-12 07:43 |
| 475d3f67-fd79-47… | Gate 3 | formula_repro |
How does the generalization of scaled tabular models trained on Criteo data perform on unseen high-cardinality categorical features in other…
COUNTEREXAMPLE HUNTER: 7.3/10 · CITATION AUDITOR: 4.5/10 · REPLICATION ATTACKER: 6.5/10
|
6.1/10 | 2026-06-12 07:43 |
| 790e88e1-86e1-4d… | Gate 3 | formula_repro |
How does the domain gap between synthetic and real-world video data affect the zero-shot accuracy of CLIP-based video encoders in gesture re…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 1.0/10
|
6.3/10 | 2026-06-12 07:43 |
| b22d1b2d-fd2b-41… | Gate 2 | unknown | How do TabPFN, CTGAN, and CausalMixFT perform in cross-domain tabular data generation tasks when evaluated on both synthetic and real-world … | - | 2026-06-12 05:07 |
| 7e2cde64-adf0-4b… | Gate 3 | formula_repro |
Can causal synthetic data generation improve the robustness of tabular foundation models against distribution shifts in cross-domain evaluat…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 7.5/10
|
8.7/10 | 2026-06-12 01:36 |
| bc4f6f71-a74f-4b… | Gate 3 | formula_repro |
To what extent does the choice of Structural Causal Model (SCM) backbone (e.g., linear vs. nonlinear) in CausalMixFT affect few-shot accurac…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
5.7/10 | 2026-06-12 01:35 |
| 41ba449b-d600-44… | Gate 3 | formula_repro |
How does the CMAL framework's image-text alignment performance on COCO and Flickr30K compare to CLIP and ALBEF in terms of Recall@1 and NDCG…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-12 01:35 |
| aa621143-3436-4a… | Gate 3 | formula_repro |
Does the scaling behavior of XSimGCL's contrastive loss formulation yield superior convergence rates compared to LightGCL when trained on de…
COUNTEREXAMPLE HUNTER: 10.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 10.0/10
|
9.8/10 | 2026-06-12 01:34 |
| 7958ccbd-1a8f-47… | Gate 3 | formula_repro |
What is the impact of the novel web-crawled data collection strategy in WebFAQ 2.0 on the domain generalization capabilities of multilingual…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-11 19:28 |
| 9f94de5f-bbb3-45… | Gate 3 | formula_repro |
What is the impact of varying the ratio of synthetic-to-real samples in CausalMixFT on the calibration error and generalization performance …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.2/10
|
9.1/10 | 2026-06-11 19:28 |
| aa6a06f7-1784-40… | Gate 3 | formula_repro |
What is the effect of curriculum learning strategies on the accuracy of large multimodal models evaluated on the MedQA benchmark?
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-11 19:28 |
| 8018fed0-9d06-45… | Gate 3 | formula_repro |
How does curriculum-based multi-task learning impact the inference latency of large multimodal models on sparse medical image-text pairs?
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-11 19:25 |
| 3c14268f-b85e-4f… | Gate 3 | formula_repro |
What is the comparative memory footprint and inference latency of multi-task trained vision-language models versus single-task baselines on …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 0.0/10
|
6.2/10 | 2026-06-11 19:25 |
| a324f8d3-15a0-49… | Gate 3 | formula_repro |
How does the stochastic inference technique in bitwise neural networks compare to other ensemble methods (e.g., snapshot ensembles, Monte Ca…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.5/10 | 2026-06-11 19:25 |
| 3bb4313b-0a73-4d… | Gate 3 | formula_repro |
To what extent does training dense retrievers on the bilingual aligned QA pairs in WebFAQ 2.0 improve alignment metrics and retrieval robust…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 7.5/10
|
8.7/10 | 2026-06-11 13:19 |
| 46e484e7-8529-4e… | Gate 3 | formula_repro |
To what extent does the inclusion of 47 million non-English WebFAQ pairs improve the robustness of multilingual encoders against domain shif…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 9.5/10
|
6.2/10 | 2026-06-11 13:18 |
| 72716d26-4a67-4e… | Gate 3 | formula_repro |
How do multilingual dense retrievers trained on SWIM-IR perform on low-resource languages in BEIR compared to models trained on natural mult…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 2.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.3/10 | 2026-06-11 13:18 |
| a9865db0-e4a2-4f… | Gate 3 | formula_repro |
How do different alignment strategies in multimodal models impact inference throughput in low-resource settings when evaluated on BRATS with…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-11 13:17 |
| e733595e-23e1-4e… | Gate 3 | formula_repro |
What is the comparative robustness of multimodal reasoning in language models with different alignment strategies when applied to cross-doma…
COUNTEREXAMPLE HUNTER: 8.2/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.7/10 | 2026-06-11 13:17 |
| 6124f4e5-dc19-47… | Gate 3 | formula_repro |
To what extent does layer-wise KV cache reconstruction in methods like ReST-KV artificially inflate needle-in-a-haystack scores relative to …
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-11 07:12 |
| 30fba3f7-6edd-44… | Gate 3 | formula_repro |
Reproducibility meta-analysis: 3 independent publications report divergent Qwen2.5 performance on Docvqa with a 80.3 percentage-point spread…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.5/10 | 2026-06-11 07:11 |
| 0595bd4f-0470-40… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.5/10 | 2026-06-11 01:23 |
| 50a4f525-1f3f-43… | Gate 3 | formula_repro |
What is the performance degradation of Unified-IO 2 on the VQA-v2 dataset when audio modalities are introduced as distractors versus text-on…
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 5.0/10 · REPLICATION ATTACKER: 7.5/10
|
6.7/10 | 2026-06-11 01:22 |
| f82ac2f4-1a92-4e… | Gate 3 | formula_repro |
How does GRACE's quantization-aware training scale with model size, and how does it affect performance on the MME and MM1K benchmarks when a…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.2/10 · REPLICATION ATTACKER: 2.0/10
|
5.9/10 | 2026-06-11 01:22 |
| cc2d0e37-a950-4a… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 3.5/10
|
7.0/10 | 2026-06-10 19:35 |
| 673590a7-25e9-41… | Gate 3 | formula_repro |
How does Qwen3's performance on GPQA Diamond compare to other frontier models when evaluated under chain-of-thought prompting versus standar…
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 9.5/10
|
6.2/10 | 2026-06-10 19:35 |
| b472d355-87a8-45… | Gate 3 | formula_repro |
How do language models compare to human experts on professional knowledge and science benchmarks v19
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.7/10 | 2026-06-10 19:34 |
| b5058ffc-3f4d-46… | Gate 3 | formula_repro |
What is the impact of million-token context windows on multimodal reasoning accuracy in Gemini 1.5 Pro versus prior versions?
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
8.8/10 | 2026-06-10 19:34 |
| 09acaf30-ab81-49… | Gate 3 | formula_repro |
To what extent does chain-of-thought prompting mitigate performance degradation in long-horizon reasoning tasks for LLMs evaluated on the Bi…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-10 19:34 |
| 031cd03f-2fbe-4d… | Gate 3 | formula_repro |
What are the benchmark performance scores of GLM-4.5-Air on reasoning mathematics coding and language understanding tasks
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 0.0/10 · REPLICATION ATTACKER: 8.5/10
|
5.8/10 | 2026-06-10 19:34 |
| 99e0cc2f-ae34-40… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.2/10 · REPLICATION ATTACKER: 8.5/10
|
8.9/10 | 2026-06-10 16:52 |
| 73ec2b2b-e67b-47… | Gate 3 | formula_repro |
What is the cross-domain generalization capability of OpenPangu-7B-MLA on empathetic speech understanding tasks when evaluated on MMSU and o…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-10 16:51 |
| dd13f070-1013-42… | Gate 3 | formula_repro |
How does the performance of self-supervised foundation models on tabular data classification compare to standard normalization techniques wh…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.0/10 | 2026-06-10 16:51 |
| 2483aaac-7f84-4c… | Gate 3 | formula_repro |
To what extent does fine-tuning on adversarial multi-hop QA examples improve the robustness of RAG systems against distractor contexts compa…
COUNTEREXAMPLE HUNTER: 9.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.5/10 | 2026-06-10 16:51 |
| 3cbe8120-1209-45… | Gate 3 | formula_repro |
How does fine-tuning on AdvRACE affect the cross-lingual robustness of MRC models when evaluated on adversarial perturbations in non-English…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-10 16:51 |
| 4a909146-446f-4d… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 2.5/10
|
6.3/10 | 2026-06-10 10:48 |
| 426ccfd7-06e6-40… | Gate 3 | formula_repro |
How does the integration of non-lexical vocal cues in multimodal language models like OpenPangu-7B-MLA affect downstream task performance on…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 3.5/10 · REPLICATION ATTACKER: 7.5/10
|
6.5/10 | 2026-06-10 10:47 |
| 294a5d5b-f300-40… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
5.7/10 | 2026-06-10 08:45 |
| 18e28019-37fb-4c… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 9.2/10
|
5.6/10 | 2026-06-10 08:45 |
| a80b4a8e-8700-4c… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-06-10 08:44 |
| e26d33b4-a5b3-48… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 6.5/10 · REPLICATION ATTACKER: 3.0/10
|
6.2/10 | 2026-06-10 08:44 |
| 30bd9c9a-90c8-4e… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
8.9/10 | 2026-06-10 08:43 |
| 3b783c5d-ec77-4e… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 9.2/10
|
5.9/10 | 2026-06-10 08:42 |
| 388b9655-1a81-4e… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 0.0/10
|
5.8/10 | 2026-06-10 08:42 |
| 42863d1d-2f6a-41… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.7/10 | 2026-06-10 08:42 |
| 0e47786d-3f42-43… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 7.5/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.2/10 | 2026-06-10 08:41 |
| 3904006d-6cfc-42… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.7/10 | 2026-06-10 08:41 |
| 845a22c0-61ad-4e… | Gate 3 | arithmetic_repro |
-
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 8.5/10 · REPLICATION ATTACKER: 8.5/10
|
8.7/10 | 2026-06-10 08:41 |
| 42a5d013-2da3-4d… | Gate 3 | unknown |
-
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.2/10 | 2026-06-10 08:36 |
| 8520660f-c1c4-4c… | Gate 3 | unknown |
-
COUNTEREXAMPLE HUNTER: 0.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
6.3/10 | 2026-06-10 08:36 |
| fa1dffe8-f9a9-4f… | Gate 3 | formula_repro |
How does the F1-score of diffusion-based tabular generative models compare to CTGAN when augmenting data for training LLMs on imbalanced tex…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-10 08:35 |
| ee851b65-000d-44… | Gate 3 | formula_repro |
What is the impact of varying the pretraining dataset size and diversity on the cross-domain generalization capabilities of tabular foundati…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-10 08:35 |
| 11c29061-cf3e-4b… | Gate 3 | formula_repro |
Does scaling the size of domain-specific training data for RAG models improve alignment with human evaluators when measured by RAGalyst's me…
COUNTEREXAMPLE HUNTER: 8.5/10 · CITATION AUDITOR: 7.5/10 · REPLICATION ATTACKER: 7.5/10
|
7.8/10 | 2026-06-10 08:35 |
| 9f6b0926-918c-40… | Gate 3 | formula_repro |
How does the scaling of unlabeled video-audio pretraining data affect the few-shot adaptation accuracy of latent action models on the RoboBe…
COUNTEREXAMPLE HUNTER: 9.0/10 · CITATION AUDITOR: 9.5/10 · REPLICATION ATTACKER: 9.5/10
|
9.3/10 | 2026-06-10 08:35 |
Math Counterexample Kills (178 total, showing 100)
Conjectures generated by the autonomous math research pipeline and killed at Gate 1 when a numerical counterexample was found. These never reach the Lean 4 proof stage.
| Conjecture ID | Problem | Statement (falsified) | Killed (UTC) |
|---|---|---|---|
| 341ec39b31764e0d… | Goldbach conjecture — computational extension | For every even integer n >= 14, there exists a Goldbach partition n = p + q (with p <= q) such that the smaller prime p lies in the interval [n/2 - sqrt(n) * ln(ln(n)), n/2]. This conjecture asserts that Goldbach partiti… | 2026-07-20 12:22 |
| 2329816d326f4d73… | Primes of form n^2+1 — density and distribution | Conjecture: For the sequence of primes of the form n^2+1, let P_k = n_k^2+1 be the k-th such prime. The difference between consecutive roots, d_k = n_{k+1} - n_k, satisfies d_k <= floor(sqrt(2)*sqrt(n_k)) for all k >= 2.… | 2026-07-20 12:12 |
| fc9fc487c02e46c8… | Twin prime conjecture — density analysis | For any integer N >= 100, let T(N) be the count of twin prime pairs (p, p+2) with p <= N. Let S_3(N) be the count of such pairs where the smaller prime p satisfies p mod 3 == 1. The conjecture states that the deviation o… | 2026-07-19 12:23 |
| 7653356c80a04bff… | Square Minus Square Factoring | For every natural number n less than 100, the square of n is either even or odd. | 2026-07-17 21:49 |
| dd33f57004a346d0… | Gauss Sum Identity | For every natural number n, twice the sum of integers from 0 to n equals n times (n+1). Specifically verified for n=100. | 2026-07-14 10:53 |
| 4628e88475724c9a… | Gauss Sum Identity | For any natural number n, twice the sum of integers from 0 to n equals n times (n+1). Specifically verified for n=100. | 2026-07-14 10:52 |
| 8b00ff270a8b4a29… | Square Minus Square Factoring | For any natural number n less than 100, the expression (n+1)^2 - n^2 is equal to 2n + 1. | 2026-07-14 10:51 |
| 2cc59bf234364fc9… | Square Minus Square Factoring | For any natural number n less than 100, the square of n is either even or odd. | 2026-07-14 10:51 |
| 1f3a254de2564028… | Goldbach conjecture — extend computational verific | For every even integer n >= 12, there exists a Goldbach partition n = p + q (with p <= q) such that the smaller prime p satisfies p < sqrt(n) * ln(n) AND p is a quadratic residue modulo the smallest prime factor of n/2. … | 2026-07-14 01:32 |
| 2f7bfd822fe44db6… | OEIS A001065 — perfect number conjecture | For every even perfect number n > 6, if n is expressed in the Euclidean form n = 2^(p-1) * (2^p - 1) where p is prime, then the sum of the decimal digits of the Mersenne prime factor M_p = 2^p - 1 is strictly greater tha… | 2026-07-13 21:21 |
| 9826a775d0564794… | Collatz conjecture — structural pattern search | For every integer n > 1, let S(n) be the set of odd integers encountered in the Collatz trajectory of n before reaching 1 (excluding the final 1). Let M(n) be the maximum element in S(n). If S(n) is non-empty, then M(n) … | 2026-07-13 21:19 |
| 2b0e9c2136934913… | Ramsey R(5,5) — improve upper bound below 48 | In any 2-coloring of the edges of K_43 that contains no monochromatic K_5, there exists no vertex v such that the red degree of v is exactly 22 AND the red neighborhood of v induces a subgraph with fewer than 130 red edg… | 2026-07-13 11:38 |
| 695e319f6af9452e… | Primes of form n^2+1 — density conjecture | For any integer n >= 2, let P_n be the set of primes of the form k^2+1 where 1 <= k <= n. Let G_n be the maximum gap between consecutive elements in the sorted sequence P_n (defining the first gap as p_1 - 1). Then G_n i… | 2026-07-13 06:50 |
| e3768564d1d14d53… | Twin prime density — Hardy-Littlewood conjecture v | For all integers x >= 10,000, the ratio of the actual count of twin prime pairs up to x to the Hardy-Littlewood prediction (2*C2*x/ln(x)^2) is strictly bounded between 0.92 and 1.08. Furthermore, the relative error |actu… | 2026-07-13 02:22 |
| f4e2bc52910d4da3… | Twin prime density — Hardy-Littlewood conjecture v | For all integers x >= 10,000, the relative error between the actual count of twin prime pairs up to x and the Hardy-Littlewood approximation (2*C2*x/ln(x)^2) is strictly bounded by 1.8 / ln(x). Specifically, |pi_2(x) - 2… | 2026-07-13 02:21 |
| e592fef2b61141e9… | Twin prime density — Hardy-Littlewood conjecture v | For all x >= 10^4, let pi_2(x) be the count of twin prime pairs up to x, and let L(x) = 2*C2*x/(ln x)^2 be the Hardy-Littlewood prediction. Define the relative error E(x) = (pi_2(x) - L(x)) / L(x). The conjecture states … | 2026-07-13 02:20 |
| 2d0ba4abd21f4366… | Ramsey R(4,6) — computational bounds | In any 2-coloring of the edges of K_35 that contains no red K_4 and no blue K_6, the red subgraph cannot contain a vertex of degree exactly 8. Specifically, the set of red degrees in such an extremal coloring must exclud… | 2026-07-12 22:16 |
| 44e2f3e458014f4f… | Twin prime conjecture — density analysis | For all integers N >= 100, let T(N) be the count of twin prime pairs (p, p+2) with p <= N. Let P_3(N) be the count of 'prime triplets' of the form (p, p+2, p+6) with p <= N. The conjecture states that the ratio R(N) = T(… | 2026-07-12 22:14 |
| ad6e319a126640d5… | Catalan's conjecture (Mihailescu) — Lean4 formal p | For any integer n > 1, if n is a perfect power (n = x^a with x > 1, a > 1), then there exists no other perfect power m = y^b (with y > 1, b > 1) in the interval (n, n + n^(2/3)], except for the specific case where n = 8 … | 2026-07-12 09:44 |
| 5806a167262a45c7… | Geometric Sum Identity | For any natural number n up to 100, the sum of the first n+1 powers of 2 (from 2^0 to 2^n) equals 2^(n+1) - 1. | 2026-07-12 09:42 |
| ddc732deacf94c22… | Sum of Odd Numbers Identity | The sum of the first 85 odd positive integers equals 85 squared. | 2026-07-12 05:36 |
| 104e5516d53b4222… | Zarankiewicz z(n,n;3,3) — improve upper bound | For n ≥ 6, the maximum number of 1s in an n×n 0-1 matrix with no 3×3 all-ones submatrix satisfies z(n,n;3,3) ≤ (n^2 - n)/2 + 1 | 2026-07-09 23:24 |
| 13761122992e44b8… | Cap set problem — F_3^n maximum | For n=6, the maximum size of a cap set in F_3^n is exactly 112. Furthermore, every maximal cap set of this size contains a subset of 28 points that forms a disjoint union of 4 affine planes of dimension 2 (2-flats), wher… | 2026-07-09 14:28 |
| 5907a60071f54a66… | Catalan's conjecture (Mihailescu) — Lean4 formal p | For any integer n > 1, if n is a perfect power (n = x^a with x>1, a>1) and the next consecutive perfect power m > n satisfies m - n = 1, then n must be 8. Furthermore, for any perfect power n > 8, the gap to the next per… | 2026-07-09 09:53 |
| 098f253a2d1544e8… | Goldbach conjecture — extend computational verific | For every even integer n > 6, there exists a Goldbach partition n = p + q (with p <= q) such that the smaller prime p lies in the interval [n/2 - sqrt(n)*ln(ln(n)), n/2]. Furthermore, the number of such 'central' Goldbac… | 2026-07-09 09:36 |
| 9d66e6a3528f42fa… | OEIS A001065 — perfect number conjecture | For any even perfect number n, the sum of its proper divisors (excluding itself) is equal to n, and the number of proper divisors (excluding itself) is greater than or equal to log(n) - 1. | 2026-07-09 02:26 |
| 7383e59f829146ba… | OEIS A001065 — perfect number conjecture | For all even perfect numbers n greater than 6, the sum of the proper divisors (excluding 1 and n) is strictly less than the geometric mean of the first and last prime factors of n. | 2026-07-09 02:26 |
| f4f07b4593f14898… | Primes of form n^2+1 — density conjecture | For every integer n >= 2, the count of primes of the form k^2 + 1 with 1 <= k <= n is strictly greater than the count of integers k in the same range such that k^2 + 1 is a product of exactly two distinct primes (semipri… | 2026-07-08 11:28 |
| b8d0385a8fb44535… | Fibonacci primes — density conjecture | For every Fibonacci prime F_p with index p > 3, the quantity (p-1)/2 is a prime number. In other words, the indices of Fibonacci primes greater than 2 are either 4 or twice a prime (Sophie Germain prime structure in the … | 2026-07-07 14:36 |
| c4be5b7578b34165… | Square Minus Square Factoring | For every natural number n less than 100, the square of n is congruent to either 0 or 1 modulo 2. | 2026-07-06 20:22 |
| 68af45652da74778… | Square Minus Square Factoring | For every natural number n less than 100, the square of n modulo 2 is either 0 or 1. | 2026-07-06 20:21 |
| c4d08438bafa482a… | Square Minus Square Factoring | For every natural number n less than 100, the square of n modulo 2 is either 0 or 1. | 2026-07-06 20:19 |
| f49d86141b87418c… | Consecutive Product Parity | For every natural number n less than 100, the product of n and its successor is divisible by 6 if and only if n is congruent to 0 or 2 modulo 3. | 2026-07-06 11:41 |
| c43b8568b87c42f0… | Goldbach conjecture — extend computational verific | For every even integer n >= 10,000, there exists a Goldbach partition n = p + q (with p <= q) such that the prime p lies in the interval [n/2 - sqrt(n)*ln(n)/4, n/2]. Furthermore, the smallest such prime p_min(n) satisfi… | 2026-07-06 11:39 |
| 57f25624e4e846c3… | OEIS A001065 — perfect number conjecture | For every even perfect number n > 6, let p be the associated Mersenne prime exponent (such that n = 2^(p-1) * (2^p - 1)). The sum of the proper divisors of the Mersenne prime M_p = 2^p - 1 is exactly equal to the integer… | 2026-07-06 07:32 |
| 61dd10cdec4c49e1… | Collatz conjecture — structural pattern search | For any integer n > 2 that is not a power of 2, let S(n) be the set of odd integers encountered in the Collatz trajectory of n before reaching 1. The product of all elements in S(n), multiplied by 3 raised to the count o… | 2026-07-06 07:31 |
| b1266844d47447cf… | Collatz conjecture — structural pattern search | For any integer n > 1, let S(n) be the set of all odd numbers encountered in the Collatz trajectory of n before reaching 1. Let M(n) be the maximum element in S(n). The conjecture states that for all n > 1 where the traj… | 2026-07-06 07:30 |
| f1945224e49f44ef… | Ramsey R(5,5) — upper bound improvement | For any edge 2-coloring of K_47 that avoids monochromatic K_5, the maximum number of monochromatic triangles is at least 4872 | 2026-07-05 11:09 |
| afc4df5791ed4a35… | Twin prime conjecture — density analysis | The number of twin prime pairs (p, p+2) with p ≤ n satisfies π₂(n) ≥ (2 * C₂ * n / (ln(n) - 1.5)) - (0.5 * sqrt(n) * ln(n)) for n ≥ 100, where C₂ is the Hardy-Littlewood twin prime constant (approximately 0.660161...). | 2026-07-04 21:00 |
| 974d71d6d011446e… | Geometric Sum Identity | The sum of the first n odd powers of 3 starting from 3^1 equals (3*(3^n - 1))/(3 - 1) for all natural numbers n. | 2026-07-03 22:39 |
| 3c35c81490be4f34… | Catalan's conjecture (Mihailescu) — Lean4 formal p | For every integer n > 1 that is not a perfect power, the open interval (n, n + c * n^(1/2) * log(n)) contains at least one perfect power x^a (with x > 1, a > 2), where c = 1.5. This conjecture refines the gap distributio… | 2026-06-30 07:36 |
| 1e9469408f494df4… | Fibonacci primes — density conjecture | For any index n > 4, if the Fibonacci number F_n is prime, then n must be a prime number p such that p is congruent to 1 or 9 modulo 10, OR p itself is equal to 3 or 5. In other words, no Fibonacci prime exists at a prim… | 2026-06-27 05:16 |
| bd796ece79f24281… | Catalan's conjecture (Mihailescu) — Lean4 formal p | For any integer n > 1, if n is a perfect power (n = x^a with x, a > 1), then the gap to the next larger perfect power m (where m = y^b > n) satisfies m - n > n^(0.45). The only exception to this bound is the pair (8, 9),… | 2026-06-27 00:35 |
| f406229dff9c4acc… | Sum of Odd Numbers Identity | The sum of the first n odd positive integers equals n squared, verified for n up to 100 using decide. | 2026-06-26 20:13 |
| 6c6df7920de647e5… | Sum of Odd Numbers Identity | The sum of the first 101 odd positive integers equals 101 squared. | 2026-06-26 20:09 |
| 072c72847c5646ea… | Square Minus Square Factoring | For every natural number n from 0 to 99, the square of n is either even or odd. | 2026-06-26 15:52 |
| b163164fc830436e… | Square Minus Square Factoring | For every natural number n less than 100, the square of n modulo 2 is either 0 or 1. | 2026-06-26 15:51 |
| 57e15775a3bc4bcc… | OEIS A001065 — perfect number conjecture | For every even perfect number n > 6, let p be the associated Mersenne prime exponent (such that n = 2^(p-1) * (2^p - 1)). The sum of the decimal digits of n, denoted S(n), satisfies the inequality S(n) < p^2 - p + 10. Th… | 2026-06-26 03:34 |
| 8ecfaac20a8c469c… | OEIS A001065 — perfect number conjecture | For every even perfect number n > 6, the sum of the decimal digits of n, denoted S(n), satisfies the congruence S(n) ≡ 1 (mod 9). Furthermore, the iterative digital root of n is exactly 1. | 2026-06-26 03:34 |
| debd6060104c4a7c… | Goldbach conjecture — computational extension | For every even integer n >= 1000, there exists a Goldbach partition n = p + q (with p <= q) such that the smaller prime p satisfies: sqrt(n) - 0.5 * ln(n) < p < sqrt(n) + 0.5 * ln(n). This conjecture claims that for suff… | 2026-06-25 23:27 |
| 81b8ac61ac714120… | Primes of form n^2+1 — density and distribution | For the sequence of primes of the form n^2+1, let p_k be the k-th such prime and g_k = p_{k+1} - p_k be the gap between consecutive primes. The conjecture states that for all k >= 2, the gap g_k is strictly bounded by 2 … | 2026-06-25 23:25 |
| e1e2e9c685dd4919… | Primes of form n^2+1 — density and distribution | For any integer N >= 100, let S_N be the set of primes p <= N such that p = n^2 + 1 for some integer n. Let M_N be the maximum gap between consecutive elements in S_N (sorted). Then M_N < 4 * sqrt(N) * ln(ln(N)). | 2026-06-25 23:24 |
| ae9f6c86bd69407f… | Primes of form n^2+1 — density conjecture | For any integer n > 100, the number of primes of the form k^2+1 with k <= n is strictly greater than the count of integers m <= n such that m^2+1 is a product of exactly two distinct primes (semiprimes). | 2026-06-25 15:06 |
| d2ab9fa0d3df4c24… | Twin prime density — Hardy-Littlewood conjecture v | For all x >= 1000, the cumulative count of twin prime pairs pi_2(x) strictly exceeds the first-order Hardy-Littlewood approximation H_1(x) = 2*C2*x/(ln x)^2, but remains bounded above by the second-order correction H_2(x… | 2026-06-25 10:41 |
| fc6cdcc2aa7a47ec… | Twin prime density — Hardy-Littlewood conjecture v | For x >= 100, the relative error between the actual count of twin prime pairs up to x and the Hardy-Littlewood prediction (2*C2*x/ln(x)^2) is strictly bounded by 1.5 / ln(x). Specifically, |pi_2(x) - 2*C2*x/ln(x)^2| / (2… | 2026-06-25 10:39 |
| 60540ad087334425… | Twin prime conjecture — density analysis | For every integer N >= 100, the count of twin prime pairs (p, p+2) with p <= N is strictly greater than the count of 'cousin' prime pairs (p, p+4) with p <= N, provided we exclude pairs where p is part of a prime triplet… | 2026-06-25 06:32 |
| 3e50957e3a5a4fd2… | Geometric Sum Identity | The sum of powers of 2 from i=0 to i=100 equals 2^101 - 1. | 2026-06-24 18:01 |
| db2c1bc8f03e4296… | Sum of Odd Numbers Identity | The sum of the first 100 odd positive integers equals 100 squared. | 2026-06-24 13:46 |
| 6e4c5386d7564a2b… | Gauss Sum Identity | For any natural number n, twice the sum of integers from 0 to n equals n times (n+1). Specifically, we verify this identity for n=42. | 2026-06-24 09:31 |
| 4ccd3991881540f6… | Square Minus Square Factoring | For every natural number n less than 100, the square of n is congruent to either 0 or 1 modulo 2. | 2026-06-24 09:27 |
| f979faa363284e15… | Square Minus Square Factoring | For every natural number n less than 100, the square of n is congruent to either 0 or 1 modulo 2. | 2026-06-24 09:26 |
| 7e2fcabce1a64039… | Quadratic Residue mod 3 | For every natural number n less than 100, the square of n modulo 3 is either 0 or 1. | 2026-06-24 05:23 |
| 70cf7e82279f45e3… | Collatz conjecture — structural pattern search | For any integer n > 1, let S(n) be the set of odd integers encountered in the Collatz trajectory of n (excluding the final 1 if the sequence terminates). Let m = min(S(n)). Then, the number of elements in S(n) that are s… | 2026-06-23 20:59 |
| d1098ee67ee042eb… | Goldbach conjecture — computational extension | For every even integer n > 6, there exists a Goldbach partition (p, q) with p + q = n such that both primes p and q can be expressed as the sum of two integer squares (i.e., p ≡ 1 (mod 4) or p=2, and similarly for q). | 2026-06-23 16:26 |
| 2ac347b7b85c4e66… | Primes of form n^2+1 — density and distribution | For the sequence of primes of the form p = n^2 + 1, let S(x) be the sum of the generating indices n for all such primes up to x. The conjecture states that S(x) is asymptotically bounded below by (2/3) * x^(3/2) / ln(x) … | 2026-06-23 16:20 |
| 5f99f6095d374442… | Primes of form n^2+1 — density conjecture | For any integer N >= 100, let S_N be the set of integers n in [1, N] such that n^2+1 is prime. The number of elements in S_N that are divisible by 3 is strictly less than the number of elements in S_N that are congruent … | 2026-06-23 07:40 |
| b29d6d023b474f30… | Primes of form n^2+1 — density conjecture | For every integer n >= 2, let P_n be the set of primes of the form k^2+1 with k <= n. Let M_n be the maximum gap between consecutive elements in the sorted sequence P_n (with the first gap defined as p_1 - 2). Then M_n i… | 2026-06-23 07:40 |
| 2e0c637f5fee4ad6… | Twin prime conjecture — density analysis | For every integer N >= 100, let T(N) be the count of twin prime pairs (p, p+2) with p <= N. Let S_odd(N) be the sum of the smaller primes p in these pairs such that p ends in the digit 1, 3, or 7 (i.e., p mod 10 in {1, 3… | 2026-06-22 23:25 |
| 79c37ffd88474fe2… | Fibonacci primes — density conjecture | For every integer n >= 5 such that the nth Fibonacci number F_n is prime, the index n must be a prime number that does not divide the class number of the quadratic field Q(sqrt(5)). Furthermore, if n is a Fibonacci prime… | 2026-06-22 14:49 |
| e640858d16aa462d… | Catalan's conjecture (Mihailescu) — Lean4 formal p | For any integer n > 1, if there exists a perfect power y^b (with y > 1, b > 1) such that n < y^b < n + n^(0.55), then n cannot be a perfect power x^a (with x > 1, a > 1) unless n = 8. This refines the gap threshold for c… | 2026-06-22 10:39 |
| 7eccfd7ac5184531… | Geometric Sum Identity | The sum of powers of 2 from 2^0 to 2^k equals 2^(k+1) - 1 for any natural number k up to 100. | 2026-06-22 10:38 |
| cb055efbafa94a5e… | Square Minus Square Factoring | For every natural number n less than 100, the square of n is either even or odd. | 2026-06-22 02:28 |
| 95a2c92217044633… | Goldbach conjecture — extend computational verific | For every even integer n >= 1000, there exists a Goldbach partition n = p + q (with p <= q) such that the smaller prime p satisfies p > sqrt(n) and the gap between the two primes is bounded by |p - q| < n^(0.45). This re… | 2026-06-21 18:17 |
| 843d975d77414a55… | Ramsey R(5,5) — upper bound improvement | In any 2-coloring of the edges of K_43 that contains no monochromatic K_5, there exists no vertex v such that the red degree of v is exactly 21 AND the red neighborhood of v induces a subgraph containing a red triangle. … | 2026-06-21 01:49 |
| e7d599128e7d45b9… | Twin prime density — Hardy-Littlewood conjecture v | For all integers x >= 100, the absolute difference between the actual count of twin prime pairs up to x and the Hardy-Littlewood prediction (2*C2*x/ln(x)^2) is strictly bounded by the square root of the prediction itself… | 2026-06-20 21:43 |
| 63dbbefc6b334b33… | Twin prime conjecture — density analysis | For every integer N >= 10,000, let T(N) be the count of twin prime pairs (p, p+2) with p <= N. Let S_odd(N) be the sum of the smaller primes p in these pairs where p ends in the digit 3 or 9, and S_even(N) be the sum whe… | 2026-06-20 17:35 |
| 342d5e78b226469c… | Fibonacci primes — density conjecture | For every index n > 4 such that the Fibonacci number F_n is prime, the index n itself must be a prime number that can be expressed as the sum of two squares (i.e., n is a Pythagorean prime or n=2). Consequently, no Fibon… | 2026-06-20 09:23 |
| cd9bb9306d6a4edf… | Fibonacci primes — density conjecture | For all integers n > 4 such that the n-th Fibonacci number F_n is prime, the index n must satisfy n ≡ 1 or 2 (mod 5). Furthermore, if n ≡ 2 (mod 5), then n must be exactly 3. Consequently, for all Fibonacci primes with i… | 2026-06-20 09:23 |
| 01cae6580a0347ec… | Geometric Sum Identity | The sum of the first n odd powers of 3 is given by the closed-form formula (3^(2n) - 1) / 8. | 2026-06-20 03:58 |
| 75ca28388123479e… | Gauss Sum Identity | For any natural number n, the sum of integers from 0 to n multiplied by 2 equals n times (n+1). Specifically verified for n=100. | 2026-06-19 18:21 |
| afd87b814d86453c… | Square Minus Square Factoring | For any natural number n less than 100, the square of n is either even or odd. | 2026-06-19 18:20 |
| 0b56f01b498740fa… | Square Minus Square Factoring | For every natural number n less than 100, the square of n is either even or odd. | 2026-06-19 18:20 |
| 7a77efd70f1a4118… | Quadratic Residue mod 3 | For every natural number n less than or equal to 100, the square of n modulo 3 is either 0 or 1. | 2026-06-19 14:15 |
| a7f5bd3a2fcc4f22… | Primes of form n^2+1 — density and distribution | For the sequence of primes of the form p = n^2 + 1, let S(x) be the set of such primes less than or equal to x. Define the 'quadratic gap ratio' for a prime p = n^2 + 1 (where n > 1) as R(p) = (p_next - p) / (2n), where … | 2026-06-19 01:53 |
| b94993ff3d514132… | Ramsey R(5,5) — upper bound improvement | In any 2-coloring of the edges of K_43 (the current lower bound for R(5,5)) that contains no monochromatic K_5, the maximum number of monochromatic K_4 subgraphs is exactly 204. Furthermore, any such extremal coloring mu… | 2026-06-18 17:32 |
| 68b38389e7ec452c… | Catalan's conjecture (Mihailescu) — Lean4 formal p | For any integer n > 1, if n is a perfect power (n = x^a with x > 1, a > 1), then the distance to the nearest other perfect power m (m != n, m = y^b with y > 1, b > 1) satisfies |n - m| > sqrt(n) * (ln(n))^0.8, with the s… | 2026-06-17 20:27 |
| 6516345988494423… | Catalan's conjecture (Mihailescu) — Lean4 formal p | For any integer n > 8 that is a perfect power (i.e., n = x^a with x, a > 1), the open interval (n, n + n^(5/6)) contains no other perfect powers. This conjecture asserts that for perfect powers greater than 8, the gap to… | 2026-06-17 20:26 |
| 2849c8ac1ec74318… | Geometric Sum Identity | The sum of the first 101 powers of 2 (from 2^0 to 2^100) equals 2^101 - 1. | 2026-06-17 20:26 |
| afc717278ab246f9… | Sum of Odd Numbers Identity | The sum of the first 42 odd positive integers equals 42 squared. | 2026-06-17 16:11 |
| 42ae3e2e624a4a21… | Square Minus Square Factoring | For any natural number n less than 100, the square of n is either even or odd. | 2026-06-17 12:05 |
| df9723f85369422d… | Square Minus Square Factoring | For every natural number n less than 100, the square of n is either even or odd (specifically, n squared modulo 2 is either 0 or 1). | 2026-06-17 12:05 |
| 849d0bcdc5b04211… | Quadratic Residue mod 4 | For every natural number n less than 100, the square of n modulo 4 is either 0 or 1. | 2026-06-17 08:01 |
| 3ab4d13b11594410… | OEIS A001065 — perfect number conjecture | For any even perfect number n > 6, let p be the largest prime factor of n (which is also the Mersenne prime exponent's base, i.e., n = 2^(p-1)*(2^p - 1)). The sum of the proper divisors of the Mersenne prime component (2… | 2026-06-16 23:52 |
| 39276ea98e4f49b0… | Primes of form n^2+1 — density conjecture | For every integer N >= 2, let P_N be the set of primes of the form k^2+1 less than or equal to N. Let M_N be the maximum gap between consecutive elements in the sorted sequence P_N (defining the first gap as p_1 - 0). Th… | 2026-06-16 07:15 |
| 816e34ad26774d21… | Twin prime density — Hardy-Littlewood conjecture v | The ratio of the actual count of twin prime pairs up to x to the Hardy-Littlewood prediction (2*C2*x/ln(x)^2) exhibits a systematic negative bias that decays according to a specific logarithmic correction term. Specifica… | 2026-06-16 03:05 |
| 1f35272d16e64f29… | Twin prime conjecture — density analysis | For any integer N >= 100, let T(N) be the count of twin prime pairs (p, p+2) with p <= N. Let S_3(N) be the count of such pairs where the smaller prime p satisfies p mod 3 = 1. The conjecture states that the deviation of… | 2026-06-15 22:59 |
| 6625e04ce42645a3… | Fibonacci primes — density conjecture | For all integers n > 3, if the n-th Fibonacci number F_n is prime, then n must be a prime number p such that p ≡ 1 (mod 4) or p = 3. In other words, there are no Fibonacci primes with prime indices p where p ≡ 3 (mod 4) … | 2026-06-15 14:49 |
| 6168a219a5854c3f… | Collatz conjecture — structural pattern search | For any integer n > 1, let S(n) be the set of odd integers encountered in the Collatz trajectory of n before reaching 1. Define the 'Odd-Step Parity Signature' P(n) as the sum of the indices (0-based) of all odd elements… | 2026-06-15 06:32 |
| d457ed3a611f478f… | Collatz conjecture — structural pattern search | For any integer n > 1, let S(n) be the set of odd integers encountered in the Collatz trajectory of n before reaching 1 (excluding the final 1). The conjecture states that the sum of the reciprocals of the elements in S(… | 2026-06-15 06:31 |
| b76082e436db434d… | Goldbach conjecture — computational extension | For every even integer n >= 10,000, there exists a Goldbach partition n = p + q (where p and q are prime) such that both p and q lie within the interval [n/2 - sqrt(n), n/2 + sqrt(n)] AND at least one of the primes p or … | 2026-06-15 02:26 |