Image AI just got nine times faster. HART, a hybrid image model from MIT and NVIDIA, produces visuals roughly nine times faster than prior diffusion systems and can run on ordinary laptops and phones. Presented at the International Conference on Learning Representations, HART pairs a fast autoregressive backbone with a compact diffusion refiner to cut compute needs. MIT also launched a Generative AI Impact Consortium to study how generative systems reshape work, hiring and design.

Faster image models, lower compute cost

Researchers at MIT and NVIDIA described a hybrid approach to image generation that combines two established techniques. The model, called HART for hybrid autoregressive transformer, uses an autoregressive component to sketch the broad layout of an image and then applies a compact diffusion module to refine details. That split lets the system reach image quality on par with state-of-the-art diffusion models while running about nine times faster.

The speed and efficiency gains matter because diffusion models typically require tens of iterative denoising steps and heavy computation. HART reduces that load, the authors say, and can run locally on a commercial laptop or smartphone. The lead authors include Haotian Tang SM ’22, PhD ’25 and Yecheng Wu, with Song Han, associate professor in the MIT Department of Electrical Engineering and Computer Science and a distinguished scientist at NVIDIA, listed as senior author. The team plans to present the work at the International Conference on Learning Representations.

Technically, yes — but the practical consequences matter. Faster generation with lower hardware needs expands who can use advanced image synthesis. It opens up real-time workflows for designers and lowers a barrier for teams that previously needed cloud GPUs or data-center time.

Where industry meets policy at MIT

MIT has also moved to coordinate academic and corporate responses to generative AI. The institute launched a Generative AI Impact Consortium that brings together university researchers and industry partners to study how generative models will change practice and policy.

Anantha Chandrakasan, dean of the School of Engineering and MIT’s chief innovation and strategy officer, leads the effort.

Chandrakasan framed the consortium’s work around the idea that generative systems and large language models are reshaping many sectors. “Generative AI and large language models [LLMs] are reshaping everything, with applications stretching across diverse sectors,” he said. The consortium zeroes in on three central questions about how humans and machines might collaborate, how behaviors around AI should be governed, and how cross-disciplinary research can yield safer tools.

Tim Kraska, associate professor in MIT’s Computer Science and Artificial Intelligence Laboratory and co-faculty director of the consortium, emphasized the need for foundational design principles. “Everybody recognizes that large language models will transform entire industries, but there's no strong foundation yet around design principles,” he said. That gap is precisely why the consortium aims to pair technical advances with societal study.

Who stands to be affected

MIT’s announcements name a wide set of application areas. The HART paper points to uses such as training robots to perform real tasks, generating scenes for video games, and helping researchers build simulated environments for self-driving cars to test rare hazards. The consortium statement highlights broader domains, from code generation to hiring processes.

Faster, cheaper generative tools will change day-to-day work for people in several roles. Examples include:

  • Designers: faster iteration and near-real-time experimentation.
  • Simulation teams: richer training worlds without massive cloud bills.
  • Recruiters and HR: expanded use of automated screening or candidate-matching aids.
  • Engineering teams: workloads that once required cloud GPUs can increasingly run locally on commodity hardware.

Related Articles

"If you are painting a landscape, and you just paint the entire canvas once, it might not look very good. But if you paint the big picture and then refine the image with smaller brush strokes, your painting could look a lot better. That's the basic idea with HART," said Haotian Tang SM ’22, PhD ’25. The team plans to present the work at the International Conference on Learning Representations.

This article was created with AI assistance.