TensorX
返回文献探索

Paper · arXiv 2408.09574

PhysBERT: A Text Embedding Model for Physics Scientific Literature

Thorsten Hellert, João Montenegro, Andrea Pollastro

8 upvotesAugust 18, 2024arXiv 预印本
AI 摘要

PhysBERT is a physics-specific text embedding model pre-trained on arXiv physics papers, surpassing general-purpose models in physics-related tasks and fine-tuning for subdomains.

text embedding modelNLPPhysBERTarXiv physics papersfine-tuningphysics-specific tasks

Abstract

The specialized language and complex concepts in physics pose significant challenges for information extraction through Natural Language Processing (NLP). Central to effective NLP applications is the text embedding model, which converts text into dense vector representations for efficient information retrieval and semantic analysis. In this work, we introduce PhysBERT, the first physics-specific text embedding model. Pre-trained on a curated corpus of 1.2 million arXiv physics papers and fine-tuned with supervised data, PhysBERT outperforms leading general-purpose models on physics-specific tasks including the effectiveness in fine-tuning for specific physics subdomains.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号