TensorX
返回文献探索

Paper · arXiv 2305.11364

Visualizing Linguistic Diversity of Text Datasets Synthesized by Large Language Models

Emily Reif, Minsuk Kahng, Savvas Petridis

2 upvotesMay 19, 2023arXiv 预印本
AI 摘要

LinguisticLens is an interactive tool for visualizing and analyzing syntactic diversity in datasets generated by large language models.

large language modelsfew-shot promptingbenchmarkingfine-tuningLinguisticLenssyntactic diversitylexical diversitysemantic diversityhierarchical visualization

Abstract

Large language models (LLMs) can be used to generate smaller, more refined datasets via few-shot prompting for benchmarking, fine-tuning or other use cases. However, understanding and evaluating these datasets is difficult, and the failure modes of LLM-generated data are still not well understood. Specifically, the data can be repetitive in surprising ways, not only semantically but also syntactically and lexically. We present LinguisticLens, a novel inter-active visualization tool for making sense of and analyzing syntactic diversity of LLM-generated datasets. LinguisticLens clusters text along syntactic, lexical, and semantic axes. It supports hierarchical visualization of a text dataset, allowing users to quickly scan for an overview and inspect individual examples. The live demo is available at shorturl.at/zHOUV.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号