New benchmark tests MLLMs' active visual observation
9/10
Researchers introduced ActiveVision, a benchmark to measure active observation in multimodal large language models (MLLMs). The benchmark consists of 17 tasks that require repeated visual perception. Top MLLMs, including GPT-5.5 and Claude Fable 5, performed poorly on these tasks, with the highest-scoring model solving only 10.6% of items. In contrast, human participants averaged 96.1% accuracy. The results suggest that current MLLMs lack robust active visual observation, highlighting the need for new architectures and training objectives. The benchmark's findings have implications for the development of more advanced MLLMs that can effectively integrate perception and reasoning.
Researchers study pretraining to post-training in large language models
8/10
A recent study explores the relationship between pretraining and reinforcement learning (RL) in large language models (LLMs), using chess as a controlled testbed. The researchers pretrained language models of varying sizes on human chess games, fine-tuned them on synthetic reasoning traces, and applied RL to chess puzzles. They found that post-RL performance is predictable from pretraining loss and that RL improves the model's policy in a way that depends on puzzle difficulty. The study also demonstrated that these findings transfer to other domains, such as math. The results provide insights into the pretraining-to-RL interface and the science of reasoning in LLMs.
A developer replaced a $120,000 bowling center system with a network of ESP32 microcontrollers at a cost of $1,600. The ESP32s were used to control various aspects of the bowling center, including scoring, lighting, and machinery. This project demonstrates the potential for affordable, open-source hardware to disrupt traditional industrial control systems. The use of ESP32s in this context highlights their versatility and capabilities in IoT applications.
A study found that people who received advice from AI were less accurate in their answers but more confident in their responses. The research suggests that AI advice can suppress critical thinking, leading to incorrect conclusions. This study involved participants who were given AI-generated answers to questions and then asked to provide their own responses. The results indicate that over-reliance on AI advice can be detrimental to decision-making. The study's findings have implications for the development of AI systems that provide advice or guidance to humans.
Moonshot AI has temporarily suspended new subscriptions due to high demand for its Kimi K3 model. The suspension is a result of the unexpected popularity of the Kimi K3, which has overwhelmed the company's capacity. This development indicates a significant interest in Moonshot AI's offerings and may reflect the growing demand for AI services. The company's response to this surge will be crucial in maintaining user satisfaction and expanding its customer base.
Anthropic uses Claude Code for large-scale code migrations
6/10
Anthropic has successfully utilized Claude Code for large-scale code migrations. This involves using AI to automatically convert and update codebases, potentially saving time and reducing errors. The technology behind Claude Code enables it to understand and transform code efficiently, which could be beneficial for companies dealing with legacy code or undergoing significant software updates. This development is relevant to the field of artificial intelligence as it demonstrates the application of AI in software development and maintenance. The use of AI in code migration can improve the efficiency and accuracy of the process.
Protester criticizes Amazon CTO over AI use in Gaza
6/10
A protester has called out Amazon's CTO for allowing Israel to use Amazon's AI technology in relation to Gaza. The issue was raised on a Reddit forum, sparking discussion. The use of AI in conflict zones raises technical and ethical concerns. The incident involves Amazon, Israel, and the use of AI in a geopolitical context.
Valve reports ongoing memory crisis with rising prices
6/10
Valve has stated that the current memory crisis, affecting hardware prices, has no end in sight. This crisis is attributed to various factors, including supply chain issues and high demand. The situation is expected to worsen, leading to increased prices for hardware components. This development is significant for the tech industry, particularly for manufacturers and consumers of gaming hardware and other high-performance computing devices. The ongoing crisis may impact the production and pricing of Steam machines and other gaming systems.
OpenAI has reduced the context size of its Codex model from 372k to 272k. This change is reflected in a recent pull request on the Codex GitHub repository. The reduction in context size could potentially improve the model's efficiency and reduce computational requirements. This update may impact the development and deployment of applications that utilize the Codex model. The change is part of ongoing efforts to optimize and refine the model's performance.
OpenAI considers open-sourcing a GPT-3-like model.
8/10
Sam Altman, in an email to OpenAI's board, discussed plans to create and release a language model similar to GPT-3 that can run on consumer hardware. The goal is to discourage others from releasing similar models and make it harder for new efforts to get funded. This move is part of OpenAI's open-source strategy discussions. The email was exposed in the Musk v. Altman case in 2026. The plan aims to preempt Stability or others from releasing comparable models.
Alibaba has introduced Triton, a programming language designed for its SAIL (Self-Adaptive Intelligent Learning) framework. Triton aims to simplify the development of AI models by providing a high-level abstraction. The language is now available on GitHub, allowing developers to explore and contribute to the project. This development is significant for the AI community as it indicates Alibaba's continued investment in AI research and development. The release of Triton may also influence the development of future AI frameworks and languages.
Loopie, a powerful looped Transformer, outperforms vanilla models.
9/10
Researchers introduced Loopie, a looped Transformer model that addresses the challenge of increasing parameter count versus looping a model. Loopie consists of two Mixture-of-Experts models with 20B and 6B parameters. Extensive studies show Loopie outperforms vanilla Transformer baselines with the same compute budget. Loopie achieves strong reasoning abilities through a novel post-training pipeline and gold-medal performance at the 2025 IMO and IPhO. This development is significant in the field of natural language processing.
Qwen.ai has announced the release of Qwen 3.8 Max. The details of this release are not specified in the given content, but it is linked to the Qwen.ai homepage for more information. Qwen.ai is involved in AI-related work, but the nature of their products or services is not detailed here. The release of Qwen 3.8 Max may be significant for those following the company's developments.
Perforce charges $500 for AI-narrated training videos
4/10
Perforce is offering AI-narrated training videos for a fee of $500. The training course, titled 'P4 Helix Core User Basic', is available on their website. The use of AI narration in training videos is a notable application of AI technology in education and corporate training. This development indicates the growing adoption of AI in various industries for content creation and dissemination.
Moonshine streams PC games to devices via Moonlight
4/10
Moonshine is an open-source project that enables game streaming from a PC to any device running Moonlight. This allows for remote gameplay on various devices, leveraging the Moonlight protocol for connectivity. The project is hosted on GitHub, where users can access and contribute to the code. Moonshine's functionality relies on the Moonlight client, which supports multiple platforms. The project's technical significance lies in its application of game streaming technology.