TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Oct 30 – Nov 5, 2023
本周最热75

CodeFusion: A Pre-trained Diffusion Model for Code Generation

Mukul Singh, José Cambronero, Sumit Gulwani +3 authors

CodeFusion, a diffusion model for code generation, outperforms auto-regressive models in natural language to code tasks by iteratively refining the entire program.

diffusion code generation modeliteratively denoisingnatural language to code generationBashHF ↗arXiv ↗

47 篇论文 · 按点赞排序

05

FP8-LM: Training FP8 Large Language Models

Houwen Peng, Kan Wu, Yixuan Wei +17 authors

A new FP8 automatic mixed-precision framework for training large language models reduces memory usage and increases speed compared to BF16 and Nvidia Transformer Engine.

34FP8low-bit data formatsHF ↗arXiv ↗
07

Learning From Mistakes Makes LLM Better Reasoner

Shengnan An, Zexiong Ma, Zeqi Lin +3 authors

LeMa, a learning-from-mistakes approach, enhances LLMs' mathematical reasoning by learning from inaccurate reasoning paths corrected by GPT-4, surpassing SOTA performance on math problems.

29Large language modelsLearning from MistakesHF ↗arXiv ↗
08

CapsFusion: Rethinking Image-Text Data at Scale

Qiying Yu, Quan Sun, Xiaosong Zhang +4 authors

CapsFusion is an advanced framework that improves multimodal pretraining data by combining web-based image-text pairs and synthetic captions, leading to enhanced model performance, sample efficiency, and scalability.

27multimodal modelszero-shotHF ↗arXiv ↗
09

Idempotent Generative Network

Assaf Shocher, Amil Dravid, Yossi Gandelsman +3 authors

A new generative modeling approach uses an idempotent neural network to project data from a source distribution to a target distribution while maintaining consistency and allowing for refinement.

25idempotentneural networkHF ↗arXiv ↗
18

Does GPT-4 Pass the Turing Test?

Cameron Jones, Benjamin Bergen

GPT-4 performed better than GPT-3.5 and ELIZA in a public online Turing Test but failed to reach human-level performance, highlighting the importance of linguistic style and socio-emotional traits in human detection.

17GPT-4Turing TestHF ↗arXiv ↗
19

E3 TTS: Easy End-to-End Diffusion-based Text to Speech

Yuan Gao, Nobuyuki Morioka, Yu Zhang +1 authors

E3 TTS is a diffusion-based text-to-speech model that generates high-fidelity audio directly from text without intermediate representations or additional conditioning, supporting zero-shot tasks.

16diffusion-basedend-to-endHF ↗arXiv ↗
21

Data-Centric Financial Large Language Models

Zhixuan Chu, Huaiyu Guo, Xinyuan Zhou +9 authors

A data-centric approach using multitask prompt-based fine-tuning and abductive augmentation reasoning enables large language models to excel in financial analysis tasks with limited labeled data.

15multitask prompt-based fine-tuningabductive augmentation reasoningHF ↗arXiv ↗
22

PockEngine: Sparse and Efficient Fine-tuning in a Pocket

Ligeng Zhu, Lanxiang Hu, Ji Lin +4 authors

PockEngine is an efficient on-device learning engine that supports sparse backpropagation and graph optimizations for fine-tuning models on edge devices with varying hardware and constraints.

15sparse backpropagationsparse updateHF ↗arXiv ↗
24

Beyond U: Making Diffusion Models Faster & Lighter

Sergio Calvo-Ordonez, Jiahao Huang, Lipei Zhang +3 authors

A novel denoising network using continuous dynamical systems improves diffusion models by reducing parameters and FLOPs while enhancing convergence speed and noise robustness.

12diffusion modelsgenerative modelsHF ↗arXiv ↗
25

Text Rendering Strategies for Pixel Language Models

Jonas F. Lotz, Elizabeth Salesky, Phillip Rust +1 authors

Character bigram rendering improves performance in pixel-based language models, enabling more compact models while maintaining equivalent performance across sentence, token, and multilingual tasks.

11pixel-based language modelstext renderersHF ↗arXiv ↗
26

What's In My Big Data?

Yanai Elazar, Akshita Bhagia, Ian Magnusson +10 authors

WIMBD is a platform for analyzing large text corpora, revealing duplicates, synthetic content, and contamination in datasets used for language models.

11WIMBDlarge text corporaHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号