TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jul 29 – Aug 4, 2024
本周最热123

SAM 2: Segment Anything in Images and Videos

Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu +15 authors

Segment Anything Model 2 (SAM 2) uses a transformer architecture with streaming memory to achieve high performance in image and video segmentation, requiring fewer interactions and faster processing than previous models.

transformer architecturestreaming memoryvideo segmentationimage segmentationHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

The Llama 3 Herd of Models

Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey +530 authors

Llama 3, a multilingual and multi-modal language model with 405B parameters, achieves competitive performance across tasks including image, video, and speech recognition when integrated through a compositional approach.

119TransformermultilingualityHF ↗arXiv ↗
03

Gemma 2: Improving Open Language Models at a Practical Size

Gemma Team, Morgane Riviere, Shreya Pathak +193 authors

Gemma 2 introduces improvements in the Transformer architecture through interleaving local-global attentions and group-query attention, showcasing superior performance relative to its size.

79Transformer architectureinterleaving local-global attentionsHF ↗arXiv ↗
04

Meltemi: The first open Large Language Model for Greek

Leon Voukoutis, Dimitris Roussis, Georgios Paraskevopoulos +6 authors

Developers created Meltemi 7B, a 7 billion parameter open-source large language model for Greek, trained on a 40 billion token corpus, and enhanced with instruction-tuning for chat applications.

68Large Language ModelMistralHF ↗arXiv ↗
05

SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal Domain

Pierre Colombo, Telmo Pires, Malik Boudiaf +7 authors

Two large legal language models, SaulLM-54B and SaulLM-141B, based on the Mixtral architecture, are introduced for domain-specific adaptation in the legal sector using continued pretraining, specialized protocols, and preference alignment with synthetic data.

66Mixtral architecturelarge language modelsHF ↗arXiv ↗
10

MindSearch: Mimicking Human Minds Elicits Deep AI Searcher

Zehui Chen, Kuikun Liu, Qiuchen Wang +4 authors

MindSearch, an LLM-based multi-agent framework, improves web information seeking and integration through parallel processing and hierarchical retrieval, achieving better performance than existing solutions.

43Large Language ModelsLLMsHF ↗arXiv ↗
11

SHIC: Shape-Image Correspondences with no Keypoint Supervision

Aleksandar Shtedritski, Christian Rupprecht, Andrea Vedaldi

SHIC leverages foundation computer vision models to learn canonical surface maps without manual supervision, achieving superior results by simulating the annotation process using image-to-image correspondences and enhanced template views.

41DensePosekeypoint detectionHF ↗arXiv ↗
14

Diffusion Feedback Helps CLIP See Better

Wenxuan Wang, Quan Sun, Fan Zhang +3 authors

DIVA enhances CLIP's performance through a self-supervised diffusion process, improving visual capabilities and multimodal understanding without additional text labels.

36Contrastive Language-Image Pre-trainingCLIPHF ↗arXiv ↗
18

A Large Encoder-Decoder Family of Foundation Models For Chemical Language

Eduardo Soares, Victor Shirasuna, Emilio Vital Brazil +3 authors

A large-scale pre-trained encoder-decoder chemical language model achieves state-of-the-art performance across various tasks by leveraging a vast dataset of molecular structures, exhibiting strong few-shot learning capabilities.

32chemical language modelslarge-scale pre-trainingHF ↗arXiv ↗
19

ThinK: Thinner Key Cache by Query-Driven Pruning

Yuhui Xu, Zhanming Jie, Hanze Dong +6 authors

ThinK, a novel query-dependent KV cache pruning method, effectively reduces memory costs by over 20% in long-context scenarios without compromising the accuracy of Large Language Models (LLMs).

32Large Language Models (LLMs)computational costsHF ↗arXiv ↗
25

OmniParser for Pure Vision Based GUI Agent

Yadong Lu, Jianwei Yang, Yelong Shen +1 authors

OmniParser enhances GPT-4V's interaction with user interfaces by accurately parsing screenshots into actionable elements, significantly improving performance on key benchmarks.

24vision language modelsmultimodal modelsHF ↗arXiv ↗
30

Matting by Generation

Zhixiang Wang, Baiang Li, Jian Wang +4 authors

Latent diffusion models with pre-trained knowledge enhance image matting by generating high-quality, detailed, and photorealistic mattes.

23latent diffusion modelspre-trained knowledgeHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号