Report on Efficient Conversational Large Language Models
Zenodo (CERN European Organization for Nuclear Research)
Abstract
Deliverable 3.1 presents the current state of efficient conversational large language model design forthe CRYSTAL project, with a specific focus on spoken interaction and deployment-relevant con-straints. The report treats conversational efficiency as a systems property spanning model behaviour,speech interfaces, runtime infrastructure, and evaluation design. Across the reviewed evidence, user-perceived performance depends on the interaction between automatic speech recognition, dialoguegeneration, speech synthesis, runtime scheduling, memory management, and evaluation protocol de-sign.The analysis indicates that the literature is most developed on inference and serving optimisation,including batching, scheduling, hardware-aware deployment, and KV-cache management. However,significant gaps remain for dialogue-specific evaluation. In particular, there is still limited standard-isation for multi-turn spoken workloads that jointly reports latency, robustness, memory growth, en-ergy behaviour, and task quality under realistic traffic conditions. This gap is relevant to CRYSTALbecause the target applications, mental health support and customer assistance, require both respon-siveness and behavioural reliability.The report also shows that spoken conversational systems must be evaluated at module boundaries aswell as within models. Interface choices between ASR, LLM, and TTS strongly influence delay, errorpropagation, interruption handling, and stability across turns. Capability-oriented review of speechLLM architectures further indicates that emotion-aware and paralinguistic modelling can improveinteraction quality, but also introduces tradeoffs in controllability, robustness, and implementationcomplexity.Based on this evidence, the practical recommendation is to defer commitment to a single unified ar-chitecture, until comparative evidence is available. For a non-integrated consortium structure, CRYS-TAL is better served by a shared, instrumented modular baseline and parallel technical tracks. Thesetracks should cover interface-efficient cascades, emotion-aware speech modelling, serving and mem-ory optimisation under multi-turn load, and limited higher-risk exploration. All tracks should operateunder common interface contracts and a common measurement protocol so results remain comparableacross partners.The deliverable provides a synthesis of current evidence and a decision framework for the next phase.It consolidates what is currently known, identifies where evidence is still weak, and defines a deci-sion framework for progressing from exploratory prototypes to benchmarked and application-facingpilots. The next CRYSTAL phase should prioritise convergence based on shared protocol evidence.Technical paths should be retained when they improve conversational outcomes under transparentprotocol controls, and revised when gains are not robust across realistic scenarios.
Authors 7
-
Affiliation as printed
Intelligent Voice Ltd
-
David Thulke Aachen
Affiliation as printed
RWTH Aachen University
-
Affiliation as printed
Universidad de Granada
-
Affiliation as printed
Universidad de Granada
-
Affiliation as printed
Universidad de Granada
Cited by 0 stored of 0
No patents citing this paper on Lens.org (checked 2026-10-06).