TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

403 篇论文 · 按点赞排序

31

Large Language Models as Optimizers

Chengrun Yang, Xuezhi Wang, Yifeng Lu +4 authors

OPRO, a method using large language models to optimize tasks described in natural language, outperforms human-designed prompts on various benchmark datasets.

79derivative-based algorithmsOptimization by PROmpting (OPRO)HF ↗arXiv ↗
33

Orca 2: Teaching Small Language Models How to Reason

Arindam Mitra, Luciano Del Corro, Shweti Mahajan +12 authors

Orca 2 enhances smaller language models' reasoning abilities by teaching them diverse solution strategies, outperforming larger models on complex reasoning tasks.

78imitation learningreasoning techniquesHF ↗arXiv ↗
35

TryOnDiffusion: A Tale of Two UNets

Luyang Zhu, Dawei Yang, Tyler Zhu +5 authors

A diffusion-based architecture unifies garment detail preservation and warping for pose and shape variation in virtual try-on tasks.

75diffusion-based architectureParallel-UNetHF ↗arXiv ↗
37

Make Pixels Dance: High-Dynamic Video Generation

Yan Zeng, Guoqiang Wei, Jiani Zheng +4 authors

PixelDance, a diffusion model-based approach, generates high-dynamic videos by incorporating image instructions for first and last frames alongside text instructions, surpassing current text-to-video methods in complexity and motion.

67diffusion modelsimage instructionsHF ↗arXiv ↗
39

FreeU: Free Lunch in Diffusion U-Net

Chenyang Si, Ziqi Huang, Yuming Jiang +1 authors

A method called FreeU improves diffusion U-Net models' generation quality by re-weighting skip connections and backbone features without additional training.

66diffusion U-NetU-NetHF ↗arXiv ↗
41

QLoRA: Efficient Finetuning of Quantized LLMs

Tim Dettmers, Artidoro Pagnoni, Ari Holtzman +1 authors

QLoRA enables efficient finetuning of large language models using 4-bit quantization and Low Rank Adapters, achieving high performance with reduced memory usage.

62QLoRALow Rank AdaptersHF ↗arXiv ↗
47

DreamLLM: Synergistic Multimodal Comprehension and Creation

Runpei Dong, Chunrui Han, Yuang Peng +11 authors

DreamLLM, a framework for Multimodal Large Language Models, directly samples in the multimodal space to enhance comprehension and creation synergy, enabling free-form interleaved content generation.

60generative modelingmultimodal spaceHF ↗arXiv ↗
51

LLM360: Towards Fully Transparent Open-Source LLMs

Zhengzhong Liu, Aurick Qiao, Willie Neiswanger +25 authors

LLM360 initiative promotes full transparency and reproducibility in LLM training by open-sourcing training code, data, model checkpoints, and intermediate results.

57Large Language ModelsLLaMAHF ↗arXiv ↗
53

Llemma: An Open Language Model For Mathematics

Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster +6 authors

Llemma, a large language model pretrained on mathematical data, outperforms existing models and demonstrates tool use and formal theorem proving capabilities.

57large language modelmathematicsHF ↗arXiv ↗
54

AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model

Seungwhan Moon, Andrea Madotto, Zhaojiang Lin +10 authors

AnyMAL is a unified model that processes multiple modalities (text, image, video, audio, IMU) and generates text, achieving top performance in multimodal tasks through a pre-trained aligner and fine-tuning with diverse instructions.

56Any-Modality Augmented Language ModelAnyMALHF ↗arXiv ↗
57

Kosmos-2.5: A Multimodal Literate Model

Tengchao Lv, Yupan Huang, Jingye Chen +11 authors

Kosmos-2.5, a unified multimodal model, generates spatially-aware and structured text from text-intensive images using a Transformer architecture and task-specific prompts.

56multimodal literate modelmachine readingHF ↗arXiv ↗
58

Generative Image Dynamics

Zhengqi Li, Richard Tucker, Noah Snavely +1 authors

A frequency-coordinated diffusion sampling process is used to predict long-term motion representations for still images, enabling dynamic video creation and interactive scene manipulation.

55frequency-coordinated diffusion sampling processneural stochastic motion textureHF ↗arXiv ↗
2 / 14

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号