BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//adass2023//talk//EEHAHK
BEGIN:VTIMEZONE
TZID:US/Arizona
BEGIN:STANDARD
DTSTART:20221107T000000
TZNAME:MST
TZOFFSETFROM:-0700
TZOFFSETTO:-0700
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
SUMMARY:Experimenting with Large Language Models and vector embeddings in 
 NASA SciX - Sergi Blanco-Cuaresma
DTSTART;TZID=US/Arizona:20231107T143000
DTEND;TZID=US/Arizona:20231107T144500
DTSTAMP:20260808T185206Z
UID:pretalx-adass2023-EEHAHK@pretalx.com
DESCRIPTION:Open Source Large Language Models enable projects such as NASA
  SciX (i.e.\, NASA ADS) to think out of the box and try alternative approa
 ches for information retrieval and data augmentation\, while respecting da
 ta copyright and users’ privacy. However\, when large language models ar
 e directly prompted with questions without any context\, they are prone to
  hallucination. At NASA SciX we have developed an experiment where we crea
 ted semantic vectors for our large collection of abstracts and full-text c
 ontent\, and we designed a prompt system to ask questions using contextual
  chunks from our system. Based on a non-systematic human evaluation\, the 
 experiment shows a lower degree of hallucination and better responses when
  self-reflection is used. Further exploration is required to design new fe
 atures and data augmentation processes at NASA SciX that leverages this te
 chnology while respecting the high level of trust and quality that the pro
 ject holds.
LOCATION:Talks
URL:https://pretalx.com/adass2023/talk/EEHAHK/
END:VEVENT
END:VCALENDAR
