Revealing Hidden Model Behaviors with Task-Specific Self-Reports
arXiv:2607.03640v2 Announce Type: replace Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice…
arXiv:2607.03640v2 Announce Type: replace Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice…
arXiv:2607.26286v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as general-purpose translation systems, but their behavior is usually evaluated under a single prompt…
arXiv:2512.17351v2 Announce Type: replace Abstract: Understanding architectural differences in language models is challenging, especially at academic-scale pretraining (e.g., 1.3B parameters, 100B tokens), where results are often…
arXiv:2607.25186v1 Announce Type: new Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal,…
arXiv:2607.18946v2 Announce Type: replace Abstract: In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling…
arXiv:2502.11150v5 Announce Type: replace Abstract: Automatic methods for scoring text readability have been studied for over a century, and are widely used in research and in…
arXiv:2607.25880v1 Announce Type: cross Abstract: LLM provenance testing asks whether a suspect LLM belongs to the same lineage as a source. Existing black-box methods largely infer…
arXiv:2607.25202v1 Announce Type: new Abstract: Conversational entrainment is well-studied in monolingual and written contexts, but remains underexplored in spoken code-switching (CSW). We present a novel cross-lingual…
arXiv:2507.23541v5 Announce Type: replace Abstract: In medical scenarios, effectively retrieving external knowledge and leveraging it for rigorous logical reasoning is of significant importance. Despite their potential,…
arXiv:2509.07829v4 Announce Type: replace Abstract: Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small open…