About This Week’s Covers
This week’s cover combines two major stories: first, an image diffusion model successfully generated a playable version of Doom without a game engine. Second, a Neuralink patient is now playing Counter-Strike 2 with just his mind. The pun is: “Look Mom, no hands… or game engine!” Here’s the Ideogram prompt: Over the shoulder shot of a man watching a game of DOOM on a large TV. He is holding up his hands, the backs of which are visible. On the screen in computer graphics font that matches the style of DOOM are the words “Look Mom, no game engine!” I added the title text in Photoshop.
The rest of this week’s covers test Ideogram’s “single-shot” ability to make title cards. Rather than choose a theme, I took a single story from each category and casually asked Ideogram to generate a cover. My goal is to see just how intuitive and accessible the prompting process can be for lay-people. Here are a few of the standouts:

Executive Summaries
Amazon Turns to Anthropic’s AI for Alexa Revamp
Amazon is preparing to launch a new version of Alexa, powered by Anthropic’s Claude AI, ahead of the 2024 holiday season. The decision comes after Amazon’s in-house AI models struggled to meet performance expectations. The new “Remarkable” Alexa will use generative AI to handle more complex queries, with Amazon planning to charge $5 to $10 per month for the service, while keeping the current version free. This move reflects Amazon’s strategy to boost Alexa’s revenue potential, as previous efforts to monetize the voice assistant, particularly through shopping features, have been largely unsuccessful. The revamped Alexa aims to compete with AI advancements from companies like Microsoft and Apple. At the moment, none of the devices in my house can remember when they have a timer or alarm set. Let’s hope that will get fixed in the free version (odds are low).
reuters
Magic AI Develops New Models with 100M Token Context, Partners with Google Cloud, and Raises $320M
Magic AI has unveiled groundbreaking models capable of handling massive amounts of information at once—up to 100 million tokens, which is equivalent to 10 million lines of code or 750 novels. In AI, a “context window” refers to how much data a model can consider at one time when making decisions. Traditional models are limited by short context windows, but Magic’s new models, called LTM-2-mini, can process much larger amounts of information, which could greatly improve tasks like code generation by having all relevant code, libraries, and documentation in context. To test these models, Magic developed a new evaluation method called “HashHop.” It challenges the model to retrieve and work with random bits of data spread across a large dataset, simulating real-world tasks like code troubleshooting, where a model may need to jump between different parts of a complex codebase. In addition, Magic AI is partnering with Google Cloud to build supercomputers using NVIDIA’s powerful chips, and has raised $320 million in new funding, bringing their total to $515 million.
Magic.dev
Andrej Karpathy Embraces AI-Assisted Programming
Andrej Karpathy, renowned AI expert, co-founder of OpenAI, and former Director of AI at Tesla, shared insights on Twitter about how AI tools are transforming programming. He’s currently using VS Code Cursor with OpenAI’s GPT-3.5 (Sonnet) instead of GitHub Copilot, finding it a significant improvement. Karpathy explains that much of his coding now involves writing prompts in English, reviewing AI-generated suggestions, and making minor edits. He likens the experience to “half-coding,” where initial code fragments guide the AI, which then generates full solutions. This method saves considerable time and effort, highlighting how rapidly AI is reshaping software development. Karpathy notes that transitioning to AI-assisted coding feels like learning to program all over again but believes there’s no turning back to traditional, unassisted methods – the only option just three years ago.
Karpathy
Claude AI Introduces “Artifacts” for Collaboration Within AI Conversations
Claude.ai has launched its “Artifacts” feature for all users, now available on Free, Pro, and Team plans, including iOS and Android apps. Artifacts provide a dedicated space within Claude conversations to create, visualize, and iterate on collaborative work. These digital objects can range from code snippets, flowcharts, and graphics to interactive dashboards and prototypes. Users can easily build on existing Artifacts, remixing and improving shared projects in real-time. For Free and Pro users, Artifacts can be shared globally, while Team plan members can collaborate securely within their groups. This feature has already seen millions of uses since its preview in June, enabling project development and interactive work sessions. It’s still too esoteric for me. Not enough people I know use Claude to enable me to collaborate, but I like the idea in theory.
Anthropic
Cerebras Introduces World’s Fastest AI Solution
Cerebras has unveiled its new AI inference platform, delivering record-breaking speeds for processing large language models (LLMs). Inference refers to the process of generating outputs—such as text, answers, or predictions—from a trained model. Cerebras’ platform can generate 1,800 tokens per second for Llama3.1-8B and 450 tokens per second for Llama3.1-70B— 20x faster than NVIDIA’s GPU-based solutions. Powered by Cerebras’ Wafer Scale Engine (WSE-3), the system achieves high speeds with significantly reduced power consumption, eliminating bottlenecks by storing entire models on-chip. This enables more complex and real-time AI tasks, like enhanced reasoning and code generation, to run efficiently. With API access and competitive pricing, Cerebras opens new possibilities for developers, starting with Llama3.1 models and expanding to larger models soon.
cerebras.ai
AI Uses Diffusion Models to Run Playable DOOM Without a Game Engine
A new AI experiment called GameNGen uses diffusion models to run a playable version of the classic game DOOM—without needing a traditional game engine. By predicting the next frame based on what happened before and the player’s actions, the AI generates each frame, allowing DOOM to run at 20 frames per second. However, it heavily “memorizes” the game, training on 0.9 billion frames, which limits its ability to work with other games or create new experiences. While it’s an impressive achievement, this AI can’t build entirely new game worlds yet. The concept holds potential, though, especially in areas like self-driving cars, where huge amounts of real-world data could train AI models to simulate complex environments.
Emollick | DrJimFan | github
OpenAI and Anthropic Partner with U.S. Government for AI Research and Model Previews
OpenAI and Anthropic have signed groundbreaking agreements with the U.S. government to collaborate on research, testing, and evaluation of their AI models. The deals, announced by the U.S. Artificial Intelligence Safety Institute, mark the first such partnerships as both companies face increasing scrutiny over the ethical and safe development of AI. The collaboration will allow the Institute, part of the U.S. Department of Commerce, to access new models from both firms and provide safety evaluations before public release. This move aligns with broader efforts to regulate AI, as California prepares to vote on a major AI development bill.
Reuters | sama | arstechnica | theverge
Google Gemini Introduces Custom “Gems” and Improved Image Generation with Imagen 3
Google is rolling out two major updates for its Gemini platform. Gemini Advanced, Business, and Enterprise users can now create “Custom Gems,” personalized AI experts tailored to specific topics or tasks. These Gems can assist with anything from coding help to career advice and are designed to save time on repetitive tasks by following user instructions (similar to OpenAI’s GPTs). In addition, Google is launching its latest image model, Imagen 3, across all Gemini users. Imagen 3 offers high-quality images in various styles, and introduces safeguards to ensure safe and ethical use (aka annoying negotiations to cajole the engine into delivering what you want). I am always wary of Google’s product team, with a few exceptions. They rarely fully bake new products and they frequently abandon their work at the expense of the user. We shall see.
Blog.google
Neuralink’s Second Patient Breaks BCI Record, Advances CAD and Gaming Capabilities
Neuralink’s latest progress report highlights impressive achievements by its second human trial participant, Alex, who received his brain-computer interface (BCI) implant. Just one day post-surgery, Alex broke the world record for BCI cursor control and showcased the Link’s potential by playing Counter-Strike 2 with his mind. Beyond gaming, he used CAD software to design a 3D-printed custom mount for his Neuralink charger, demonstrating the implant’s versatility. Neuralink continues to improve performance and functionality while addressing prior technical issues, with a focus on enhancing autonomy for those with disabilities. The Neuralink episode of Lex Friedman is worth listening to. There is a lot of AI involved in the software side of the device.
Neuralink
AI Visuals and Charts
“As promised, here is a breakdown of how I did the Deadpool animation I recently posted.
“Luchador Action Figure Animation 💪 Tools I used for this: @ideogram_ai for generating reference images @ViggleAI to transform me into a Luchador @AdobeAE for compositing @ComfyUI to improve the results
““The Last Meatball” Animation: @jboogx_creative & enigmatic_e Movement: @jboogx_creative & @8bit_e Reference Imagery: FLUX on @hellocivitai Jboogx and I have known each other for the better part of the last 12 months. I think we’ve both been inspired by each other’s work and
Top 59 Links of The Week – Organized by Category
Agents and Copilots
“Generate over 10,000+ words in just one minute with LongWriter and then deploy with vllm. This is based on the recent “LongWriter” paper, where they introduced AgentWrite pipeline that decomposes ultra-long generation tasks into subtasks. It enables off-the-shelf LLMs to
Augmented and Virtual Reality (AR/VR)
“Can you convert architecture design stage from 200 to 7 days? Yes sir. First look at Gaudi-v0 model. Introducing the world first text-to-building platform. Users describe their new project – Gaudi generates an initial 3D mesh and converts it to a structural model.
“Wow. Incremental 3D reconstruction at >50 keyframes/second 🤯 This new paper Spann3R proposes having external spatial memory that you query to predict the 3D structure of the next frame in world space. Can process ordered collections in real-time!
Sketch2Scene
Automatic Generation of Interactive 3D Game Scenes from User’s Casual Sketches
“The new virtual try-on model from Kuaishou kcolors is very impressive.
“It’s time to explore the virtual ‘playground’ where AI meets DoD acquisition needs! 👉
Autonomous Vehicles
“Thrilled to announce our strategic partnership to bring @wayve_ai’s Embodied AI to @Uber’s global mobility network. 🌏 Together we’re massively ramping our AI’s fleet learning and collaborating with automotive OEMs to bring autonomous mobility to consumers sooner. Excited to
Uber will win self-driving | Manas J. Saloi
Chips, Hardware, and Infrastructure
Nvidia Earnings: Revenue Jumps 122%, a Positive Sign for AI Boom – The New York Times
Nvidia’s Future Relies on Chips That Push Technology’s Limits – WSJ
Education AI
Personalizing education with ChatGPT | OpenAI
Ethics/Legal/Security AI
California Legislature Approves A.I. Safety Bill – The New York Times
California legislature passes controversial “kill switch” AI safety bill – Ars Technica
“Don’t believe the SB 1047 hype folks! Few models thus far created would be covered (only those that cost $100M+), and their developers are voluntarily doing extensive safety testing anyway I think it’s a prudent step, but I don’t expect a huge impact either way” / X
“AI Agenda: OpenAI has shown its ‘Strawberry’ AI to national security officials. And, what could a ‘Strawberry’ product look like?
“The information reporting on OpenAI Strawberry 🍓 Their guess is it’s an advanced reasoning mode that you’ll be able to turn on/off depending on how time sensitive your queries are. They’re also reporting that OpenAI has been demoing this capability to the NatSec community.
“Being pro AI regulation doesn’t mean someone must be in favor of every AI bill on pain of hypocrisy. If you’re in favor of AI regulation, you should probably want the first iterations of it to be very good. Bad regulation can undermine the case for good regulation going forward.” / X
“This was a cool listen. I think Cloud+AI is increasingly making the @levelsio -style model of a scrappy solo serial micro-entrepreneur viable, allowing one person to spin up and run a number of companies that generate income, possibly well into billion-dollar valuations.” / X
Exclusive: Google DeepMind Staff Push to End Military Contracts | TIME
Google rolls out safeguards for more of its AI products ahead of the US presidential election | TechCrunch
Chinese firms bypass US export restrictions on AI chips using AWS cloud | CIO
China’s AI Engineers Are Secretly Accessing Banned Nvidia Chips – WSJ
X’s Grok directs to government site after sharing false election info – The Verge
‘Being on camera is no longer sensible’: persecuted Venezuelan journalists turn to AI | Venezuela | The Guardian
Google AI
“Chatbot Arena update⚡! The latest Gemini (Pro/Flash/Flash-9b) results are now live, with over 20K community votes! Highlights: – New Gemini-1.5-Flash (0827) makes a huge leap, climbing from #23 to #6 overall! – New Gemini-1.5-Pro (0827) shows strong gains in coding, math over
“Today, we are rolling out three experimental models: – A new smaller variant, Gemini 1.5 Flash-8B – A stronger Gemini 1.5 Pro model (better on coding & complex prompts) – A significantly improved Gemini 1.5 Flash model Try them on
“Google just released three new experimental Gemini 1.5 models: – A smaller 8B parameter Flash model – An updated Pro model – An improved Flash model The new 1.5 Pro now ranks as #2, and the new 1.5 Flash ranks as #6 on the Chatbot Arena leaderboard!
“Google researchers just developed GameNGen, an AI that can simulate DOOM at over 20 fps It works by predicting each frame in real time with a diffusion model At scale this could mean AI will be able to create games on the fly, personalized to each player
“We just shipped a new native prompt gallery in Google AI Studio ✨ Test out long context, native multi-modal (image, video and audio), structured outputs, and more!
Google appoints former Character.AI founder as co-lead of its AI models | Reuters
International AI
China’s robot makers chase Tesla to deliver humanoid workers | Reuters
Russia Could Take Out West’s Internet, No Good Back up Plan – Business Insider
China’s AI Engineers Are Secretly Accessing Banned Nvidia Chips – WSJ
Meta AI
Meta leads open-source AI boom, Llama downloads surge 10x year-over-year | VentureBeat
With 10x growth since 2023, Llama is the leading engine of AI innovation
“Open source AI is the way forward and today we’re sharing a snapshot of how that’s going with the adoption and use of Llama models. Read the full update here ➡️
Multimodality
“Microsoft’s new open source Phi 3.5 vision model is really good at OCR/text extraction — even on handwriting! You can prompt it to extract tabular data as well. It’s permissively licensed (MIT). Play around with it here:
“InterTrack: tracking human-object interaction without any pre-scanned object template. We build the shapes from scratch and track them from a single RGB video. Key idea: 4D tracking = one global shape + per-frame pose. Great collaboration with @GerardPonsMoll1 @janericlenssen
OpenAI
OpenAI says ChatGPT usage has doubled in the last year
OpenAI Races to Launch ‘Strawberry’ Reasoning AI to Boost Chatbot Business — The Information
“ORION – OAI’s Giant Leap Forward – The Plan Is to Leave Competition Far Behind 😱 The Information reports that OAI’s all-powerful reasoning model, Strawberry, is available internally but has not yet been released. Instead of launching Strawberry, they are using it to produce” / X
Exclusive | Apple, Nvidia Are in Talks to Invest in OpenAI – WSJ
Apple and Nvidia may invest in OpenAI – The Verge
Podcasts/YouTube/Op-Eds
“Struggling to keep up with AI’s breakneck pace? I’m launching a quick daily AI news digest. I sift through tons of AI sources every day. I’ve been sharing a roundup on our @huggingface Slack, and @ClementDelangue suggested making it public. Smart idea! Here’s what you’ll get: -” / X
“Here’s my conversation with Pieter Levels (@levelsio), self-taught developer and entrepreneur who designed, programmed, shipped, and ran over 40 startups, many of which are hugely successful. In most cases, he did it all by himself, while living the digital nomad life in over 40
“Grace Hopper is a national treasure and so is this video recently declassified by the NSA. She is full of wisdom that we all can learn from.
Oprah Winfrey to Host ‘AI and the Future Us’ Special for ABC
Publishing
“Generate over 10,000+ words in just one minute with LongWriter and then deploy with vllm. This is based on the recent “LongWriter” paper, where they introduced AgentWrite pipeline that decomposes ultra-long generation tasks into subtasks. It enables off-the-shelf LLMs to
“You can Crawl entire website with Claude 3.5 or GPT4-o with this open-sourced tool firecrawl. 💯 Turn entire websites into LLM-ready markdown or structured data. Scrape, crawl and extract with a single API. Crawls all accessible subpages and give you clean data for each.
Google AI Overviews rollout ramps up, hitting publisher visibility
(RAG) Retrieval-Augmented Generation
“Long Context vs. RAG Paper Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach This research paper, conducted by Google DeepMind, provides guidance on whether it’s better to use Long Context natively or leverage RAG, specifically
Robotics and Embodiment
“Real Jets, Real Autonomy. In a groundbreaking series of tests this July, Shield AI, alongside partners @KratosDefense and @parrylabs , demonstrated the future of air combat. High-performance, autonomously controlled MQM-178 Firejet drones flew alongside human-piloted jets in
“Chinese electric car manufacturer XPENG teased the 2nd generation of its humanoid robot It’s set to be fully revealed at the company’s 1024 Tech Day on October 24th China’s robotic developments/news in the past two weeks have been insanely impressive
“A new paper by Disney Research introduces a 2-stage technique for motion tracking in physics-based character animation, using a pretrained latent space and RL to control diverse and unseen motions on both virtual and real-world humanoid robots. Paper:
“We chant in English (prompting), and silicon bricks learn to control a set of metal (Optimus) to manipulate other sets of atoms. Inanimate components are being brought to life by invisible tokens from a neural network. If this isn’t magic I don’t know what is.” / X
Science and Medicine
World-first lung cancer vaccine trials launched across seven countries | Lung cancer | The Guardian
AI spots cancer and viral infections with nanoscale precision
Customizable AI tool developed at Stanford Medicine helps pathologists identify diseased cells | News Center | Stanford Medicine
Google and Others Are Developing AI That Can Hear Signs of Sickness – Bloomberg





Leave a Reply