BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//pyladiescon-2024//talk//FN3RLV
BEGIN:VEVENT
UID:pretalx-pyladiescon-2024-FN3RLV@pretalx.com
DTSTART:20241207T051000Z
DTEND:20241207T053000Z
DESCRIPTION:n the rapidly evolving field of natural language processing\, t
 he evaluation of large language models (LLMs) is crucial for understanding
  their performance and guiding their development. This talk delves into th
 e two primary evaluation methodologies: reference-based and referenceless 
 techniques.\n\nReference-based evaluation relies on predefined ground trut
 h references to assess the quality of generated text. Metrics such as BLEU
 \, ROUGE\, and METEOR are commonly used to compare the generated output ag
 ainst these references\, providing insights into the model’s accuracy an
 d fluency. However\, these metrics often fall short in capturing the nuanc
 es of human language and creativity.\n\nOn the other hand\, referenceless 
 evaluation techniques\, such as BERTScore\, perplexity\, and human judgmen
 t\, offer a complementary perspective by assessing the coherence\, relevan
 ce\, and overall quality of the generated text without relying on referenc
 e texts. These methods can better capture the subtleties of language gener
 ation and provide a more holistic view of model performance.\n\nThis talk 
 will explore the strengths and limitations of both evaluation approaches\,
  highlighting recent advancements and practical applications. This talk is
  suitable for anyone who has basic understanding of NLP and LLMs and want 
 to know more about evaluation stratergies
DTSTAMP:20260713T180834Z
LOCATION:Main Stream
SUMMARY:Holistic Evaluation of Large Language Models: From References to Hu
 man Judgment - Riya Joshi
URL:https://pretalx.com/pyladiescon-2024/talk/FN3RLV/
END:VEVENT
END:VCALENDAR
