OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
Tested on Semianalysis’s InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.
Read original article ↗
Related Articles
Ox Alpha is a multimodal GLM model by Z.ai: Apple Mini and Studio desktops target AI power users
Perplexity moves its computer use agent to on-device models. OpenAI’s Jalapeño may be faster than Nvidia’s best chips. I
Learning never stops: How AI makes learning continuous
OpenAI’s new report explores how students and educators use ChatGPT to make learning more continuous, with support that
Instead of using ChatGPT to write the articles, the operators used the chatbot to strip any wording that might mark them