FACTOR evaluates language model factuality by transforming a factual corpus into a benchmark of true vs. similar incorrect statements, demonstrating improved accuracy over perplexity in large models and with retrieval augmentation.
Yinghao Xu, Wang Yifan, Alexander W. Bergman +3 authors
Layered surface volumes (LSVs) enhance 3D GANs by efficiently representing articulated digital humans with fine details and high fidelity from unstructured 2D images.
Abhirut Gupta, Ananya B. Sai, Richard Sproat +5 authors
A method for generating phonetically corrupted language text from non-native speakers improves benchmarks and highlights the need for more robust language models.
T2I-CompBench is a benchmark for assessing compositional text-to-image generation, introducing evaluation metrics and GORS, a reward-driven fine-tuning method.
7text-to-image modelsGenerative mOdel fine-tuning with Reward-driven Sample selectionHF ↗arXiv ↗
Online test-time training (TTT) framework for video frames improves performance by using a small temporal window compared to fixed-model or offline TTT methods.
M2C, a morphologically-aware framework, tests and highlights generalization failures of NLP models to specific typological characteristics across 12 diverse languages.
Speech-LLaMA integrates acoustic information into text-based LLMs using Connectionist Temporal Classification and a simple audio encoder, demonstrating improved performance on multilingual speech-to-text tasks using a decoder-only architecture.
A new dataset of 3.2 million visual instruction tuning pairs improves multimodal performance in visual perception, reasoning, and planning by training on high-quality, diverse manual annotations.
7foundation modelslarge language modelsHF ↗arXiv ↗
The study investigates the generalization challenges in imitation learning for visual robotic manipulation by quantifying the impact of different factors of variation in both simulation and real-world settings.
Arvind Mahankali, Tatsunori B. Hashimoto, Tengyu Ma
Transformers with linear self-attention trained on synthetic linear regression tasks learn to implement gradient descent, with distributional changes affecting preconditioned gradient descent or nonlinear functions.
Fabian Paischer, Thomas Adler, Markus Hofmarcher +1 authors
The paper presents novel methods for constructing semantic mappings between image and language embedding spaces, enabling effective image captioning without extensive gradient information.
6pretrained language modelsmulti-modal modelsHF ↗arXiv ↗
Adam Fisch, Amal Rannen-Triki, Razvan Pascanu +4 authors
A benchmark of task sequences for continual learning evaluates transfer scenarios, and a selective initialization strategy for new models is proposed to mitigate negative transfer.
Lewis Ho, Joslyn Barnhart, Robert Trager +8 authors
International institutions may have an important role to play in ensuring
advanced AI systems benefit humanity. International collaborations can unlock
AI's ability to further sustainable development, and coordination of regulatory
efforts can reduce obstacles to innovation and the spread of benefits.
Conversely, the potential dangerous capabilities of powerful and
general-purpose AI systems create global externalities in their development and
deployment, and international efforts to further responsible AI practices could
help manage the risks they pose. This paper identifies a set of governance
functions that could be performed at an international level to address these
challenges, ranging from supporting access to frontier AI systems to setting
international safety standards. It groups these functions into four
institutional models that exhibit internal synergies and have precedents in
existing organizations: 1) a Commission on Frontier AI that facilitates expert
consensus on opportunities and risks from advanced AI, 2) an Advanced AI
Governance Organization that sets international standards to manage global
threats from advanced models, supports their implementation, and possibly
monitors compliance with a future governance regime, 3) a Frontier AI
Collaborative that promotes access to cutting-edge AI, and 4) an AI Safety
Project that brings together leading researchers and engineers to further AI
safety research. We explore the utility of these models and identify open
questions about their viability.
5Commission on Frontier AIAdvanced AI Governance OrganizationHF ↗arXiv ↗
Solvent is a unified protein folding framework that supports various state-of-the-art models, enabling consistent and fair comparisons in the protein structure modeling field.
A new reinforcement learning framework using multi-granularity unit test feedback enhances code generation with large language models, achieving superior performance on benchmarks.
5reinforcement learninglarge language modelsHF ↗arXiv ↗
Markus Anderljung, Joslyn Barnhart, Jade Leung +21 authors
Advanced AI models hold the promise of tremendous benefits for humanity, but
society needs to proactively manage the accompanying risks. In this paper, we
focus on what we term "frontier AI" models: highly capable foundation models
that could possess dangerous capabilities sufficient to pose severe risks to
public safety. Frontier AI models pose a distinct regulatory challenge:
dangerous capabilities can arise unexpectedly; it is difficult to robustly
prevent a deployed model from being misused; and, it is difficult to stop a
model's capabilities from proliferating broadly. To address these challenges,
at least three building blocks for the regulation of frontier models are
needed: (1) standard-setting processes to identify appropriate requirements for
frontier AI developers, (2) registration and reporting requirements to provide
regulators with visibility into frontier AI development processes, and (3)
mechanisms to ensure compliance with safety standards for the development and
deployment of frontier AI models. Industry self-regulation is an important
first step. However, wider societal discussions and government intervention
will be needed to create standards and to ensure compliance with them. We
consider several options to this end, including granting enforcement powers to
supervisory authorities and licensure regimes for frontier AI models. Finally,
we propose an initial set of safety standards. These include conducting
pre-deployment risk assessments; external scrutiny of model behavior; using
risk assessments to inform deployment decisions; and monitoring and responding
to new information about model capabilities and uses post-deployment. We hope
this discussion contributes to the broader conversation on how to balance
public safety risks and innovation benefits from advances at the frontier of AI
development.
A framework combining large language models, visual-language models, and model-based planning synthesizes robust robot trajectories for diverse manipulation tasks from natural language instructions.
4large language modelsrobot manipulationHF ↗arXiv ↗
Anthony Simeonov, Ankit Goyal, Lucas Manuelli +5 authors
The proposed system, trained from demonstrations, rearranges objects in 3D scenes by iteratively refining poses and focusing on relevant local geometric features to handle multi-modality and generalization.
43D point cloudsiterative pose de-noisingHF ↗arXiv ↗
Belinda Z. Li, Jason Eisner, Adam Pauls +1 authors
A study explores real-time spoken interruption for dictation and editing using large pre-trained language models, showing a trade-off between accuracy and latency.
4large pre-trained language modelsend-state accuracyHF ↗arXiv ↗
Combining narrow robotic imitation datasets with broad human video demonstrations improves generalization of eye-in-hand visuomotor policies using partial observability and image masking.