a junk drawer in a kitchen is labeled “Nerd Stuff”
“Introducing The Together Enterprise Platform Run faster inference, fine-tuning, and training on your own models, in any environment (including your own VPC or on-prem). Our customers have achieved 2-3x faster inference and reduced costs by up to 50%, on their own infrastructure.
“This paper in Nature caught my attention. Suggests that larger and more instructable LLMs may become less reliable. It investigates LLMs across three elements: difficulty concordance, task avoidance, and prompting stability. “We also find that early models often avoid user
[2409.16211] MaskBit: Embedding-free Image Generation via Bit Tokens
[2409.14720v1] ControlEdit: A MultiModal Local Clothing Image Editing Method
[2409.17778v1] Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEs
Liquid AI raises $37.6M to build ‘liquid’ neural networks – SiliconANGLE
“PyTorch has historically been super slow for a lot of “small” RL workloads, but the reason is kinda dumb. RL often uses tiny neural networks, and end up ridiculously CPU overhead bound. Luckily, with cudagraphs and torch.compile, we often see >5x speedups! I remember the first” / X
DeepSeek-Coder-V2/paper.pdf at main · deepseek-ai/DeepSeek-Coder-V2
[2409.17093v1] BitQ: Tailoring Block Floating Point Precision for Improved DNN Efficiency on Resource-Constrained Devices
[2409.16159v1] ComiCap: A VLMs pipeline for dense captioning of Comic Panels
Exploring Parallel Strategies with Jax | AstraBlog
2409.13598
MMMU
[2409.14939v1] FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
[2409.02060v1] OLMoE: Open Mixture-of-Experts Language Models
Commit-0
“Very proud to see this acceptance especially given widespread adoption of Elo in NLP benchmarking. Special shoutout to @mziizm who championed this work, and to @mellem_boo who led this work while part of our scholars program. 📜 Paper link:
paper.pdf
“Should I drop a wallpaper app ?
[2409.17703v1] PGN: The RNN’s New Successor is Effective for Long-Range Time Series Forecasting





Leave a Reply