X · @AndrewYNg
· X / Twitter
New course: Build LLM applications that respond to user requests quickly by running on hardware designed for fast inference. This short course was bui…
New course: Build LLM applications that respond to user requests quickly by running on hardware designed for fast inference. This short course was built with @Cerebras and taught by @zhennydez, @duerr_seb, and @MilksandMatcha.When a model generates text, much of the time is spent moving its weights out of memory and in