Image created with Flux Pro Ultra. Image prompt: A Minecraft screenshot featuring a blocky school building with bookshelves, crafting tables set up as desks, and pixelated educational posters on walls, with “EDUCATION” written in pixelated Minecraft font across the top

Agentica – Home https://agentica-project.com/

“Do regular AI detoxes — where you try to write that two-pager without any AI intervention, where you make that decision without asking four different LLMs, where you create that piece of art from scratch — just so we don’t lose touch with what really makes us human. Our” / X https://x.com/bilawalsidhu/status/1908780786947367413

“Play Pals: Built using Lovable AI Some issues exist, but Lovable credit is depleted. Still, it’s fun to build. https://x.com/BugNinza/status/1906053712298361078

“@lennysan Really excited that PMs are learning about evals – incase it might be interesting we are doing a deep dive into the subject for engineers in this course We will cover all the nitty gritty of “how” including agents, multi-turn and several edge cases https://x.com/HamelHusain/status/1910163448757150076

“we are cooked https://x.com/svpino/status/1910506102002753941

“We’re open-sourcing BrowseComp (“Browsing Competition”), a new, challenging benchmark designed to test how well AI agents can browse the internet to find hard-to-locate information. It’s like an online scavenger hunt…but for browsing agents. https://x.com/OpenAI/status/1910393421652520967

“Google Gemini 2.5 is the first public AI model to definitively beat the level of performance of human PhDs with access to Google on hard multiple choice problems inside their field of expertise (around 81%). All AI tests are flawed, but GPQA Diamond has been a pretty good one.” / X https://x.com/emollick/status/1907737487176286418

Announcing Google’s 2025 Growth Academy: AI for Health cohort https://blog.google/outreach-initiatives/entrepreneurs/growth-academy-ai-health-2025/

“I’ve been saying that DeepSeek will expand from verifiable to general domains, and expected a paper. Here is that paper. Self-Principled Critique Tuning. rule-based online RL. Gemma-2 27b is enough to match R1. This is roughly what Google does for Gemma 3 and likely Geminis. https://x.com/teortaxesTex/status/1907987423377666538

“New Anthropic research: How university students use Claude. We ran a privacy-preserving analysis of a million education-related conversations with Claude to produce our first Education Report. https://x.com/AnthropicAI/status/1909626720476365171

“Which degrees have the most disproportionate use of Claude? Perhaps not surprisingly, Computer Science leads the field, with 38.6% of Claude conversations related to the subject, which makes up only 5.4% of US degrees. https://x.com/AnthropicAI/status/1909626726612717942

“Anthropic released Claude for Education focused on developing students’ critical thinking It has a new “Learning Mode” that guides students through problem-solving rather than giving straight up answers to their questions! https://x.com/adcock_brett/status/1908913597884874973

“Look who we found hanging out in her new @StanfordEng Gates Computer Science office! We’re truly delighted to welcome @YejinChoinka as a new @stanfordnlp faculty member, starting full-time in September. ❤️ https://x.com/stanfordnlp/status/1908178010127397005

“Stanford students’ research discussion forum AlphaXiv introduced Deep Research for arXiv The tool compiles literature reviews from trending papers, turning hours of research work into mere seconds of natural language search https://x.com/rowancheung/status/1909845159748976999

“chatgpt plus is free for college students in the US and canada through may!” / X https://x.com/sama/status/1907862982765457603

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading