a sunny day at the beach. a sand sculpture of a castle is crumbling. the letters “AGI” logo is drawn in the sand –ar 5:3 –style raw

Introducing Hard Prompts Category in Chatbot Arena | LMSYS Org

“The most common LLM failures you see shared on Twitter are word games – AI is really bad at working with text positions (“give me 10 sentences that end with the word apple” “give me 3 countries that end with a k”). This approach suggests that the problem may be solvable.” / X

“Honestly this is sort of unsettling to watch. From a YouTuber who is doing experiments with AI-powered NPCs in VR: five historical figures are on a train, only the YouTuber is human, the others are LLMs. Watch the LLMs figure out who is the dumb human. 

Elon says AGI next year.  Not sure if he’s joking.   “@OfficialLoganK Next year” / X

“We are launching SEAL Leaderboards—private, expert evaluations of leading frontier models. Our design principles: 🔒Private + Unexploitable. No overfitting on evals! 🎓Domain Expert Evals 🏆Continuously Updated w/new Data and Models Read more in 🧵 

SEAL leaderboards

Reward Bench Leaderboard – a Hugging Face Space by allenai

“NVIDIA presents NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models Achieves #1 on the MTEB leaderboard 

“A great complementary benchmark to LMSys Arena – private, clean, trusted third party evaluation of public models. Scale AI is doing an outstanding community service here. 👏” / X

“Wharton Professor Ethan Mollick explains that Artificial General Intelligence (AGI) is a pivotal moment in humanity’s trajectory. In a recent discussion, Ethan Mollick, a professor at the Wharton School and author of “Co-Intelligence: Living and Working with AI,” explored the 

Ways to think about AGI — Benedict Evans

“Was curious so added a bunch more LLM-as-a-judge models to RewardBench. GPT4 still > GPT4 turbo (april update) > GPT4o Llama 3 70b Prometheus 2 8x7b > claude 3 haiku > Prometheus 7b > GPT3.5 Llama 3 8b just behind gpt3.5 

“LLMs “intelligence” is hard to benchmark, as we don’t have good benchmarks for human performance at complex tasks. Take theory-of-mind: several tests found GPT-4 beats humans, but another one finds a huge gap. Is it the testing structure? Prompting? Which is right? Hard to know. 

“4/ While LMSYS and other efforts in the community are awesome, we still think there’s a lot to be desired in 3rd party evaluations. One of our design principles is to produce evals that are impossible to overfit. As we saw with our prior GSM1k research, we think it’s critical” / X

“Nice, a serious contender to @lmsysorg in evaluating LLMs has entered the chat. LLM evals are improving, but not so long ago their state was very bleak, with qualitative experience very often disagreeing with quantitative rankings. This is because good evals are very difficult 

“AI versus 100,000 humans in creativity in this careful study using the Divergent Association Test (a well-validated measure, but all measures of creativity have flaws) GPT-4 wins. Better prompting can further improve performance & diversity of ideas. 

“Nobody understands the full blown revolution we are about to experience in enterprise software. Build your self a little mental picture. An AI prompt window. Just sitting there ready for input. Then watch all your customer data go into that prompt. All of it Then your ERP” / X

“Transformers Can Do Arithmetic with the Right Embeddings Achieves up to 99% accuracy on 100 digit addition problems by training on only 20 digit numbers with a single GPU for one day repo: 

“Transformers can learn arithmetic with the right embeddings. Models trained on 20-digit addition can generalize to 100-digit addition. The same tricks can do 15 digital multiplication, sorting, etc. 🧮 Unlike the previous SOTA, we don’t need fancy hardware. We do training” / X

AI displacing 50% of jobs by 2027 is ‘uncannily accurate’: Kai-Fu Lee | Fortune

Bearish AGI Predictions

“Motivation: even powerful LLMs struggle to attend to concepts like sentences as they index by token. See examples of GPT4 & Llama 2 failing in the figure. This is a fundamental flaw in the architecture. How can we achieve AGI with a model that can’t do that?! 🧵(2/5) 

“Yann LeCun @ylecun: “There is no such thing as AGI and we shouldn’t be talking about AGI at all” 

“The Doomer’s Delusion: 1. AI is likely to kill us all 2. Hence AI must be monopolized by a small number of companies under tight regulatory control. 3. Hence AI systems must have a remote kill switch. 4. Hence foundation model builders must be eternally liable for bad uses of” / X

“General intelligence, artificial or natural, does not exist. Cats, dogs, humans and all animals have specialized intelligence. They have different collections of skills and an ability to acquire new ones quickly. Much of animal and human intelligence is acquired through” / X

“The emergence of superintelligence is not going to be an event. We don’t have anything close to a blueprint for super intelligent systems today. At some point, we will come up with an architecture that can take us there. The design will start by having the intelligence level of” / X

“VCs are betting against (1) continued scaling, where larger models beat specialized models while rapidly decreasing costs & (2) AGI, which would invalidate a lot of the underlying assumptions for AI startups (& many firms) It is an okay bet but I wonder if it is a conscious one” / X

“Whats interesting about this list of recently funded AI startups is that most of them are a bet that AGI will not happen in the next 5-8 years (average time to exit) I am not saying that they are wrong, but it is a strong private signal of VC beliefs contrasting with public ones” / X

“Is AI sentient? My friend and colleague Prof. John Etchemendy, a renowned professor and co-Director of @StanfordHAI , just co-authored this piece to debunk the claim that today’s LLMs are sentient @TIME https://twitter.com/drfeifei/status/1793753017701069233

Leave a Reply

Trending

Discover more from Ethan B. Holland

Subscribe now to keep reading and get access to the full archive.

Continue reading