Skip to content
X · @AndrewYNg · X / Twitter

New course: Build LLM applications that respond to user requests quickly by running on hardware designed for fast inference. This short course was bui…

New course: Build LLM applications that respond to user requests quickly by running on hardware designed for fast inference. This short course was built with @Cerebras and taught by @zhennydez, @duerr_seb, and @MilksandMatcha.When a model generates text, much of the time is spent moving its weights out of memory and in