BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//pyladiescon-2024//speaker//MTABDN
BEGIN:VEVENT
SUMMARY:Holistic Evaluation of Large Language Models: From References to H
 uman Judgment - Riya Joshi
DTSTART:20241207T051000Z
DTEND:20241207T053000Z
DTSTAMP:20260907T063712Z
UID:pretalx-pyladiescon-2024-FN3RLV@pretalx.com
DESCRIPTION:n the rapidly evolving field of natural language processing\, 
 the evaluation of large language models (LLMs) is crucial for understandin
 g their performance and guiding their development. This talk delves into t
 he two primary evaluation methodologies: reference-based and referenceless
  techniques.\n\nReference-based evaluation relies on predefined ground tru
 th references to assess the quality of generated text. Metrics such as BLE
 U\, ROUGE\, and METEOR are commonly used to compare the generated output a
 gainst these references\, providing insights into the model’s accuracy a
 nd fluency. However\, these metrics often fall short in capturing the nuan
 ces of human language and creativity.\n\nOn the other hand\, referenceless
  evaluation techniques\, such as BERTScore\, perplexity\, and human judgme
 nt\, offer a complementary perspective by assessing the coherence\, releva
 nce\, and overall quality of the generated text without relying on referen
 ce texts. These methods can better capture the subtleties of language gene
 ration and provide a more holistic view of model performance.\n\nThis talk
  will explore the strengths and limitations of both evaluation approaches\
 , highlighting recent advancements and practical applications. This talk i
 s suitable for anyone who has basic understanding of NLP and LLMs and want
  to know more about evaluation stratergies
LOCATION:Main Stream
URL:https://pretalx.com/pyladiescon-2024/talk/FN3RLV/
END:VEVENT
END:VCALENDAR
