TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 25 – Dec 31, 2023
本周最热61

SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling

Dahyun Kim, Chanjun Park, Sanghoon Kim +15 authors

A novel technique, depth up-scaling (DUS), efficiently enhances large language models (LLMs) without complex changes, building SOLAR 10.7B that outperforms existing open-source models in various NLP tasks, including instruction-following.

depth up-scalingDUSmixture-of-expertsMoEHF ↗arXiv ↗

47 篇论文 · 按点赞排序

10

Unsupervised Universal Image Segmentation

Dantong Niu, Xudong Wang, Xinyang Han +3 authors

A novel unsupervised model, U2Seg, achieves high performance in instance, semantic, and panoptic segmentation through self-supervised learning and clustering.

21unsupervised image segmentationSTEGOHF ↗arXiv ↗
11

UniRef++: Segment Every Reference Object in Spatial and Temporal Spaces

Jiannan Wu, Yi Jiang, Bin Yan +3 authors

UniRef++ unifies four reference-based object segmentation tasks using a single architecture with a UniFusion module and a unified Transformer, achieving state-of-the-art performance on RIS and RVOS and competitive results on FSS and VOS.

20referencing image segmentationfew-shot image segmentationHF ↗arXiv ↗
12

DreamGaussian4D: Generative 4D Gaussian Splatting

Jiawei Ren, Liang Pan, Jiaxiang Tang +4 authors

DreamGaussian4D efficiently generates 4D content with reduced optimization time, controllable motion, and high-quality animated meshes using Gaussian Splatting representation.

19Gaussian Splattingspatial transformationsHF ↗arXiv ↗
15

Reasons to Reject? Aligning Language Models with Judgments

Weiwen Xu, Deng Cai, Zhisong Zhang +2 authors

A novel framework, Contrastive Unlikelihood Training (CUT), effectively aligns large language models using language feedback, outperforming existing methods with less data and iterative improvements.

18large language modelslanguage feedbackHF ↗arXiv ↗
18

LangSplat: 3D Language Gaussian Splatting

Minghan Qin, Wanhua Li, Jiawei Zhou +2 authors

LangSplat constructs a 3D language field using a splatting technique with hierarchical semantics, enabling precise and efficient open-vocabulary querying in 3D spaces while significantly outperforming existing methods.

15CLIPNeRFHF ↗arXiv ↗
20

Pangu-Agent: A Fine-Tunable Generalist Agent with Structured Reasoning

Filippos Christianos, Georgios Papoudakis, Matthieu Zimmer +13 authors

A framework integrates structured reasoning into AI agents' policies using intrinsic and extrinsic functions, enhancing adaptability and performance over diverse tasks by combining prior knowledge with modular learning.

15Reinforcement Learning (RL)Large language models (LLMs)HF ↗arXiv ↗
23

YAYI 2: Multilingual Open-Source Large Language Models

Yin Luo, Qingchao Kong, Nan Xu +50 authors

YAYI 2, a multilingual large language model with 30 billion parameters, achieves superior performance in Chinese contexts after pre-training and fine-tuning compared to other open-source models.

14large language modelsLLMsHF ↗arXiv ↗
24

Exploiting Novel GPT-4 APIs

Kellin Pelrine, Mohammad Taufeeque, Michał Zając +2 authors

Exploration of gray-box threats on GPT-4 API functionalities, including fine-tuning, function calling, and knowledge retrieval, reveals new vulnerabilities that can enable harmful outputs and unauthorized actions.

13fine-tuningfunction callingHF ↗arXiv ↗
25

Human101: Training 100+FPS Human Gaussians in 100s from 1 View

Mingwei Li, Jiachen Tao, Zongxin Yang +1 authors

Human101 is a framework that reconstructs high-fidelity dynamic 3D humans from single-view videos in real-time using 3D Gaussian Splatting and Human-centric Forward Gaussian Animation, achieving high rendering speeds and quality.

123D Gaussian SplattingHuman-centric Forward Gaussian AnimationHF ↗arXiv ↗
28

Parrot Captions Teach CLIP to Spot Text

Yiqi Lin, Conghui He, Alex Jinpeng Wang +3 authors

CLIP models exhibit a text spotting bias in vision-language applications due to excessive reliance on visual text within images in datasets like LAION-2B.

11CLIPvision-language applicationsHF ↗arXiv ↗
29

InsActor: Instruction-driven Physics-based Characters

Jiawei Ren, Mingyuan Zhang, Cunjun Yu +3 authors

InsActor, a generative framework using diffusion-based human motion models, produces high-quality physics-based animations guided by human instructions through diffusion policies and compact latent skill sequences.

10diffusion-based human motion modelsdiffusion policiesHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号