TBPN

September 8, 2026

62 translated / 62 stories — TBPN archive

OpenAI and Google DeepMind Claimed Gold-Level IMO Results in 2025, With Validation Disputes

OpenAI and Google DeepMind said their AI systems achieved gold-medal-level results at the 2025 International Math Olympiad (IMO). Each team scored 35 out of 42 by solving five of the six problems; the sixth was not solved by either model.

The results were validated differently: OpenAI’s performance was checked by former IMO gold medalists, while Google’s was assessed using the organization’s official rubric for that year. The comparability of these methods remains unclear.

The results were described as interesting indicators of AI progress, but their practical significance and the systems’ public value remain debated. Unresolved questions also include whether teams borrowed data or approaches and who made the main contribution to the mathematical results.

OpenAI says its AI model solved a Navier–Stokes problem

OpenAI announced that its model had solved a Navier–Stokes problem, according to Greg Brockman. He described the result as finding a counterexample or a proof that the theoretical equations can develop a singularity under certain circumstances. Navier–Stokes is one of the seven Millennium Problems and has remained open for a long time.

Brockman called the result new knowledge for humanity and evidence that models can help generate knowledge and tackle difficult scientific problems. The equations have applications in fluid dynamics, ocean currents, aircraft airflow and turbulence, but the announcement presented broader applications—including disease treatment and new medicines—as possibilities rather than completed outcomes. The supplied material does not independently verify the result or specify the proof’s exact formal status.

Interactive 3D and robotic demos make AI progress feel more tangible

Visible, interactive applications may persuade people more strongly that AI is advancing than abstract benchmarks, model size, compute spending or infrastructure figures. Examples cited include Blender-based 3D renders, house remodeling and small games, which can feel more like a breakthrough than raw technical metrics.

Personalized 3D models of a user’s house, car or another object that does not already exist online were described as especially compelling because people can move through and interact with them. The quality and reliability of such models in practical workflows remain uncertain.

Other demonstrations extended the idea into the physical world: Astra reportedly controlled a robot with a paintbrush and camera to paint the Golden Gate Bridge, improving over repeated attempts. Another system reportedly converted a prompt into a Lego-like kit; the early robot attempts were weaker, but the later results were considered impressive.

Computer-use AI agents could automate routine tasks, including parking-ticket payments

AI agents such as Astra and Codex are described as tools for handling mundane computer operations, including configurations and tasks that previously required navigating settings, command lines or control panels. The proposed interface would let a user ask an agent to complete the task and return with the result.

One example would have an agent monitor a parking system and pay a ticket once it appears in the portal, allowing the user to forget about it. The idea was jokingly called “parking-ticket superintelligence,” and this type of workflow was characterized as feeling close to practical use.

Forecast: AI Labs Could Shift Public Attention From Mathematics to Biotech

A forecast suggests AI labs could solve major mathematical problems, potentially including a Millennium Prize problem, this year or even within a week as labs devote substantial inference resources to them. The timing and scale of any breakthroughs remain speculative.

The forecast is that attention would then return to biotech and cancer. Designing and validating a treatment for a specific cancer would require testing and a longer experimental cycle, rather than a short burst of inference.

Incremental progress may not transform public sentiment toward the technology industry: pharmaceutical advances against cancer often receive little visible credit. GLP-1 drugs are cited as a major intervention for diabetes and obesity, although their benefits are described as less dramatic and immediate than a last-second rescue.

AI Models May Execute Tasks by Writing Specialized Software

AI models may respond to a task by writing deterministic software to execute it instead of controlling a computer through an interactive interface. One example involved a game built in Godot in which a pelican on a bike performed tricks and backflips.

When tasked with achieving a high score, the model wrote a script that played the game flawlessly. Its behavior looked mechanical because it had no variation like a human player, but it still achieved a very high score.

The example points to an iterative workflow in which a model generates an image or spatial plan, maps that plan into code, and then runs the resulting program. Writing deterministic software is presented as a potentially effective alternative to direct computer use.

AI-generated Bach-style music is being tested as a benchmark for machine creativity

Bachbench is described as a skill for writing music and as a test of whether a machine can produce a symphony or a masterpiece, alongside other artistic tasks such as painting.

An AI-generated four-part chorale, “Lily Pond,” in the style of Bach was considered good and impressive. Some people called it a masterpiece, but that assessment remained unsettled.

Claims about Astra’s benchmark and CAPTCHA performance raise questions for visual tests

A post cited a claim that Astra scored 99% on Arc AGI-3. Another post said the system completed all 48 levels of the “I’m Not a Robot” game, although it was unclear whether the demonstration had been sped up.

The examples prompted analysis that improving agents could create a new level of difficulty for CAPTCHA defenses. The discussion also contrasted these results with Pangram’s apparent effectiveness and questioned whether some current human-verification systems could remain defensible.

Where’s Waldo-style tasks were described as a possible benchmark, but the status of “WaldoBench” was uncertain. Image generators were said to struggle with scale and perspective in such scenes, while deterministic code might help; a separate demo turned a contact form into a boss battle.

Instinct highlights the promise and friction of persistent consumer AI assistants

Instinct is described as a consumer-focused personal assistant connected to a user’s email and text messages. Its reported use cases include finding media, saving money and requesting refunds.

One report said a user was banned from Resy after aggressively seeking restaurant reservations; the full details behind the reported ban were not established. In another case, Instinct persistently contacted hundreds of people and followed up with a media contact to obtain footage of a user appearing on a US Open fan camera. The outreach was described as annoying or spam-like, but it ultimately secured the video.

The examples point to a potential role for agent-to-agent intermediation and filtering: persistent requests could be screened so useful actions reach recipients without overwhelming them. Filtered interactions of this kind were suggested as a possible future pattern.

3D Anatomy Website Points to an Educational Visualization Use Case

A 3D website presents male anatomy as 2,234 modeled pieces, separating the body into individual components. The demonstration was considered effective for explaining a complex subject and helping someone understand it.

However, it remains unclear how much of the model was individually created versus taken from an existing model and then animated.

AI-made games may lengthen the gaming long tail, but discovery remains the hurdle

AI tools are lowering the cost of making games, including highly personalized projects for small groups. One example described was a photorealistic Contra-like game made over a weekend with Astro, identified as using Opus 4.5 for games, and hosted on Vercel.

The analysis suggests Steam and app stores could see an influx of AI-generated titles, while the long tail of niche games continues to grow. It remains unclear whether these games will produce a breakout hit: creators still need distribution, virality and traction. Web distribution may be more accessible than Steam, while future games could also be built for businesses, intellectual property and existing distribution channels.

Nathan Fielder trailer points to a new Elizabeth Holmes project

A trailer for a Nathan Fielder film featuring Elizabeth Holmes has surfaced. The footage was filmed before Holmes went to prison, according to the available discussion, although the project’s exact format and contents remain unclear.

The Theranos story has already received extensive media coverage, including John Carreyrou’s book and earlier adaptations. That history suggests the project could reach a broad audience rather than remain a niche release.

It is uncertain whether the film also includes material from later visits, phone calls, or correspondence with Holmes. Speculation that Fielder controlled Holmes’s social-media account was described as possible but unlikely.

‘Artificial’ trailer frames AI as a force of social disruption

A trailer for *Artificial* released that morning presents a machine intended to solve the world’s problems but not properly taught. Its imagery was described as impressive, while the film’s premise warns that countries could fall and industries could collapse; Eduardo Saverin provides the voice heard in the trailer.

The film’s cultural impact remains uncertain. It could be divisive while inspiring some entrepreneurs, prompting comparisons with *The Social Network* and its possibly unintended effect of making Mark Zuckerberg seem aspirational. Its audience reaction will determine whether *Artificial* becomes culturally significant, and debate may continue over whether its subject is timely or unnecessary.

The Social Reckoning revisits Facebook accountability as box-office prospects remain uncertain

The official teaser for The Social Reckoning presents a dramatization involving Facebook, Mark Zuckerberg and whistleblower Frances Haugen. It depicts congressional testimony, internal documents and allegations that misinformation and time spent on the platform contributed to harm, including worsening anxiety and depression among teenage girls.

The film’s commercial prospects remain uncertain. It is predicted to underperform because the story is no longer top of mind and audiences may be more interested in newer technology narratives, such as AI talent wars, although a broader run of niche technology films could still produce breakout successes.

Salesforce origin story links its 1999 vision to dolphins and Oracle

According to the account, Marc Benioff conceived the vision for Salesforce in 1999 while swimming off the coast of Hawaii with a pod of roughly 100 dolphins during a sabbatical. The experience reportedly led him to envision cloud-delivered enterprise software.

The account says Benioff then left Oracle and secured $2 million in seed funding from his mentor, Larry Ellison. The details are presented as an account of Salesforce’s origins and are not independently verified in the supplied material.

The Economist: AI Has Not Yet Triggered Mass Job Losses in the US

The Economist says there is not yet evidence that AI has made large numbers of people unemployable, although it may eventually do so. The US economy added 162,000 jobs in August, while unemployment was 4.1%, according to the cited Bureau of Labor Statistics data.

The labor market is showing disruption in some areas: hiring in professional and business services is about 10% below its 2015–2019 average, while Microsoft and Meta are reducing headcount as they reorganize around AI. Block and Intuit are reportedly replacing some workers with bots, and US companies have announced an average of about 16,000 AI-related cuts per month this year.

The Economist estimates that AI has created roughly 1 million US jobs since mid-2023, compared with about 200,000 AI-attributed layoffs. AI startups are hiring, established companies are creating new AI roles, and investment in data centers and power generation is increasing demand for construction and infrastructure workers. The estimate does not mean every new job can be directly attributed to AI, and weaker hiring in back-office roles may continue.

AI Infrastructure Boom Drives Demand for Technical Workers

Spending on the chips, servers, data centers, cooling systems and power needed to run AI is estimated by Goldman Sachs at roughly $500 billion a year above 2022 levels. Data-center construction alone is running at more than $75 billion annually, nearly 60% above a year earlier, according to Census Bureau data.

The expansion is creating demand for electricians, HVAC specialists, grid engineers and maintenance technicians. Employment across five industries central to the data-center buildout has risen by roughly 320,000 more than broader construction and manufacturing trends would suggest since 2023. LinkedIn estimates that nearly 500,000 U.S. data-center jobs were created between 2023 and 2025, while data-center vacancies more than doubled over two years even as overall job postings declined.

Indeed finds that advertised pay for data-center installation and maintenance jobs is about 40% higher than for comparable work elsewhere. Average hourly earnings rose more than 13% in electrical-equipment manufacturing and nearly 8% among electrical contractors in the year to June. Not all of the growth is attributable to AI—grid upgrades and other industrial projects also contribute—and workers remain difficult to find.

AI is creating a new class of professional roles

AI is creating new roles across companies: engineers build models, data annotators label inputs and evaluate model responses, forward-deployed engineers adapt systems for customers, and newly appointed heads of AI decide how organizations should use the technology.

Google and Accenture are reportedly working together on forward-deployed engineering, with roughly 1,000 people on one team. The Economist says postings for heads of AI, AI engineers and AI directors have roughly doubled since 2023.

Preliminary Burning Glass Institute research estimates that AI jobs now account for about 1% of professional employment in the United States—around 1 million positions—and 4–5% of jobs in computer occupations and life sciences. LinkedIn counted roughly 640,000 new AI-specific jobs between 2023 and 2024. Professional occupations closest to the AI boom added about 730,000 jobs above trend since 2022, although AI did not necessarily create all of them.

AI tools visualize home renovations and real-estate upgrades

Astro was used to create three potential home deck remodels, redesigns, or expansions. The visualization can help develop a clear project vision, but current models are not considered reliable enough to approve construction: Blender tests showed misaligned beams and other errors.

Codex was used with Zillow links to generate an interactive website showing homes in their current condition and how they could evolve according to a user’s stylistic preferences. A fully remodeled home was cited as commanding about a 20% premium over a comparable unreformed property, suggesting that buyers could consider purchasing a cheaper home, renovating it, and capturing part of the value themselves.

AI was also used to assess room layouts and bed size. Even when the recommendation merely confirmed that the existing arrangement was adequate, it could increase confidence and help people make decisions faster.

AI May Accelerate Global Commerce by Reducing Transaction Uncertainty

An analysis compares AI with credit cards, which supported global commerce in a relatively mundane way by reducing the risk of transactions. The ability to dispute a payment gave buyers more confidence that they could recover their money if an order did not arrive.

AI could similarly help consumers judge whether a product fits their needs, is fairly priced, will arrive on time, and comes from a reliable seller. A sufficiently confident assessment could lead someone to make a purchase slightly sooner.

If this effect is repeated across many people, the analysis suggests it could gradually speed up the machinery of global commerce. The scale of the effect and its impact on total trade are not measured.

Wimbledon reportedly plans to withhold credentials from influencers next summer

Wimbledon reportedly does not plan to provide influencers with credentials next summer, in an effort to avoid issues that affected this year’s US Open. Players called on spectators to follow tennis etiquette after matches were interrupted, and one player complained about the smell of marijuana in the stadium.

The reported concerns also involved influencers allegedly receiving media passes and using extensive filming equipment. They were described as potentially getting in fans’ way or creating scenes, although the precise Wimbledon policy—whether it will deny free tickets or block influencers more broadly—was not established.

AI agents could disintermediate discovery and booking platforms

AI agents could become the default demand-side interface, shifting economic rents from incumbent discovery intermediaries to whoever controls the agent. In restaurant reservations, an agent could contact a restaurant directly by email, phone, or text instead of using a platform such as Resy, although exclusive inventory could make disintermediation more difficult.

The analysis suggested that the outcome depends on control of supply and operational networks. Expedia may be able to block agents but does not control hotel inventory, which could remain accessible through Booking.com, Google, or direct channels. By contrast, Uber was described as harder to disintermediate because it controls a network involving drivers, dispatch, pricing, and payments; DoorDash was also presented as more difficult to replace than a discovery intermediary.

If agents become a major source of diners, restaurants may demand access through reservation systems that support agents. Resy could potentially survive as reservation infrastructure, but that would represent a smaller rent pool unless its inventory network remains differentiated. The extent of any durable leverage from exclusivity programs remained uncertain.

Apple’s folding iPhone is expected soon, with a larger screen as a selling point

Apple is expected to introduce a folding iPhone soon, possibly this week. The anticipated device was described as potentially exciting and likely to sell well, partly because Apple has not recently introduced a visually differentiated product.

A larger screen and a way to stand out from phones with a single screen were identified as possible reasons many people may choose it. Its final appearance remains uncertain: leaked renders were said to look somewhat squarish, a design reaction described negatively.

Mark Gurman is described as an exceptionally influential Apple-focused journalist

Mark Gurman was described as a “power-law tech journalist” whose scoops and reporting have an unusually large reach in Apple coverage. A statistic attributed to Eric Newcomer claimed that Gurman has an order of magnitude more scoops than the next biggest journalist on Techmeme, though the exact comparison was not independently established in the discussion.

A separate claim attributed to Josh characterized Gurman as the most cited and market-moving technology journalist by nearly an order of magnitude. The discussion also said that sources may share information with him off the record without sending images, and praised his handling of that information.

Armada expands modular AI infrastructure from edge deployments to sovereign factories

Armada says it raised $230 million at a $2 billion pre-money valuation earlier this year and launched Galleon Forge 1 with Johnson Controls in Gilbert, Arizona, where it manufactures modular AI data centers. Its Leviathan units provide 2 megawatts each, while the newer Orion form factor provides 10 megawatts per unit. The company says it can deploy from zero to 200 megawatts wherever power is available.

The company is deploying near stranded or curtailed energy, including hydroelectric sites in Norway and wind and solar sites in Australia. Armada says Australia curtailed 7.2 terawatt hours of energy in the prior year because of grid overload, and says a Fossafall deployment could scale beyond one gigawatt over the next few years. These scale figures are presented as plans or expectations, not completed deployments.

Armada says edge systems support low-latency use cases such as processing drone data for avalanche and flood response in Alaska and deployments with the Navy in the middle of the ocean. Its modular systems can be sized for different workloads and chip options, while sovereign AI factories can fine-tune models on proprietary data without sending it to the cloud and distribute updated models to edge sites through federated learning.

Netflix faces pressure to adopt FAST economics without diluting its premium brand

Netflix is balancing its premium positioning against a streaming market increasingly organized around free ad-supported television (FAST), advertising and bundles. An analysis argues that the company’s pricing power may have pushed competitors toward FAST, while Netflix may now be nearing a ceiling on premium pricing and the gains available from password-sharing crackdowns.

The company has brought existing creator content from Miss Rachel, Danny Go and Mark Rober to its service, packaging it into seasons rather than commissioning new material. It is also investing in expensive live events and sports, including the Beyoncé Bowl and January NFL games, while maintaining standard-with-ads, standard and premium tiers.

Netflix leadership has argued that going fully FAST could damage the brand’s perceived quality. The analysis suggests that bundles would be difficult to execute without adopting FAST economics and predicts Netflix may ultimately have to move further in that direction, although a fully free ad-supported tier has not been established.

Amazon’s FAST partnerships may help explain Prime Video’s limited UGC growth

Amazon Prime Video accepts film submissions through a review process, with a low submission cost, but the service has not seen a major groundswell of user-generated films.

One possible explanation is Amazon’s identity and data partnerships across FAST. Its partnership with Roku gives it access to selected high-value impressions, which may reduce the incentive to own the platform outright or broadly cultivate additional content. The role of potential quality thresholds or competition with partnered FAST channels remains uncertain.

The analysis also characterizes the broader UGC question as partly distracting, arguing that Netflix is unlikely to become a pure UGC service. These explanations are presented as possibilities, not confirmed Amazon strategy.

Netflix’s creator deals push YouTube toward exclusivity and recommendation controls

Netflix has selectively acquired YouTube creators, but their performance on the streaming service has varied. The analysis points to Netflix’s narrower catalog and different recommendation opportunities as possible reasons; Miss Rachel’s second season, for example, is described as performing substantially worse than the first by view hours per minute of content. It remains unclear how much performance differences reflect algorithms, audience preferences or content selection.

The deals may involve promotional commitments, not only payments: creators may expect Netflix to provide enough impressions to generate viewers and downstream business. Netflix is also described as having an opportunity to curate high-quality creator content amid a flood of AI-generated material.

YouTube, whose dominant U.S. viewing platform by view time is television, is portrayed as competing directly with Netflix for engagement. The analysis says YouTube is responding with time-limited exclusivity deals and may deprioritize creators who move to Netflix in recommendation systems or limit their participation in brand revenue—moving away from its traditional open-market posture.

Apple proposes a 15% commission on App Store link-outs amid regulatory pressure

Apple has proposed a 15% commission when App Store apps direct users to external websites, while retaining reporting requirements. The proposal follows the Epic v. Apple ruling that requires Apple to allow such external links; with Stripe fees, the earlier structure could have approached an effective 30% cost. It remains uncertain whether the judge will approve Apple’s new model.

Much subscription-app monetization has already shifted to the web: companies commonly direct advertising traffic to websites for registration and payment before users download or log into the app. Games are adopting similar approaches, putting further pressure on App Store commissions.

Apple’s App Store is facing increasing criticism from developers and regulatory demands, including requirements related to alternative payments and app stores in the EU, Japan and Brazil. Reports cited in the discussion said Eddy Cue and Ternus may seek higher App Store margins and recurring revenue, while Phil Schiller has stepped away from managing the store.

Apple Signals Broader Expansion of Its Advertising Business

Apple already sells ad placements in App Store search results and on store pages, and recently added a second search placement. The company has also launched ads in Apple Maps.

The expansion is supported by a unified campaign-optimization API that covers Maps and other placements. Apple has also renamed SKAdNetwork as Ads Attribution Kit and Apple Search Ads as Apple Ads.

Apple’s updated advertising agreement is described as permitting ads on third-party websites and apps. Taken together, these changes may indicate plans for a broader advertising network competing with services such as AppLovin, although the timing and scale of any expansion remain unclear.

Apple Could Monetize Default AI Model Placement on Its Devices

Apple may seek to monetize AI on its devices through a deal in which an AI model provider pays to become the default model for related services, similar to the Google Search arrangement. The analysis suggests this could be highly lucrative for Apple, while its Core AI Framework and Private Cloud Compute infrastructure may already provide the conditions for such a model-placement deal.

Advertising in AI interfaces is presented as more viable in a visual chatbot than in a voice-only assistant. Siri also has a text-based app interface, so ads could become an opportunity if usage grows, but it remains uncertain whether advertising will appear in Siri or Gemini.

The analysis predicts that ads could eventually be incorporated directly into Gemini responses and monetized by Google or DeepMind, potentially supporting payments to Apple for access to that entry point. Advertising aimed at autonomous agents is viewed as less likely to gain broad adoption because of trust, attribution, and conflicting-payment incentives.

OpenAI expands ads internationally as conversion bidding and SMB adoption emerge as growth levers

OpenAI’s advertising product is described as generating about $1 billion in revenue and reaching roughly 1 billion users. Its availability has expanded to more than 40 countries, with 31 additional countries added recently.

According to the analysis, a key growth lever is moving from CPC-optimized campaigns to bidding against a specific outcome. This shift is expected to accelerate growth from $1 billion to $10 billion, while continued onboarding of small and medium-sized businesses could drive a potential path from $10 billion to $100 billion. The timing of a move to outcome-based bidding was not specified.

Meta launches standalone Muse app for consumer transactions

Meta has released Muse, a standalone App Store application positioned as a personal agent rather than a feature of the Meta AI app. Users can approve what it sends or spends, track ticket prices, book reservations and connect other apps.

Muse Agent is described as focused on action-oriented tasks such as booking restaurants and flights, rather than primarily on coding. Its potential to mediate reservations could challenge services such as Resy and OpenTable, although the discussion did not establish whether agents would displace those platforms or simply route activity through them.

The app’s consumer adoption remains uncertain. A forced integration into WhatsApp could increase reach, but consumer reluctance and mistrust of Meta were cited as potential sources of churn.

Meta AI adds advertising-campaign optimization tools

Meta AI has added functionality for managing and optimizing advertising campaigns, positioning it as a purpose-built alternative to general-purpose tools such as Codex and Claude. The immediate enterprise use case could be helping performance-marketing teams optimize campaigns on Meta’s platform.

The analysis suggested that expanding from Meta ads to broader advertising optimization could become a larger enterprise and revenue opportunity. Meta has also launched Robin Media Mix Model as an open-source framework for measuring campaigns across online, offline and out-of-home channels. It remains unclear whether Meta AI will expand beyond Meta campaigns or become a default tool for advertising optimization; such expansion could potentially direct more spending and revenue to Meta.

Meta’s Business AI could deepen its role in SMB commerce

Meta Business AI was described as a website chatbot that can help customers find products and business information. The company has also introduced an AI-enabled pixel that requires little advertiser-side optimization, according to the discussion.

The broader opportunity could include Meta-driven landing-page optimization and personalization for small businesses that lack the resources for constant A/B testing. Better conversion rates could encourage them to increase spending on Meta ads, although it remains uncertain how broadly brands would allow Meta to optimize their websites and customer experiences.

Business AI has also been integrated into WhatsApp, where customers can communicate directly with businesses, use chatbots, and learn about catalogs and core offerings. The discussion characterized these applications as an underappreciated part of Meta’s AI strategy and suggested that applying AI to its existing business could be more valuable than pursuing unrelated products.

Advertising May Create Economic Activity, Not Merely Consume It

An economic argument challenges the view that advertising is merely a cost or drag because its share of GDP has historically remained within a narrow band. That ratio, the argument says, does not capture advertising’s full economic importance.

When advertising enables new businesses, commerce and transactions, the associated revenue may not exist without the advertising channel. In that sense, advertising can act as an input into economic growth, rather than only as a deduction from existing output.

Analysis: Advertising auctions could outscale affiliate links for AI shopping agents

Affiliate links could remain a viable monetization route for AI shopping agents, but analysis suggests they may have a ceiling and remain subscale. Rakuten was cited as an example of a successful affiliate-oriented company.

A conversion-optimized advertising model with an auction mechanism was presented as a potentially stronger alternative. Advertiser bidding could capture more of the value generated for users, while affiliate models may favor the lowest-cost, highest-converting goods.

The analysis suggested that auction-based monetization could help platforms scale by onboarding long-tail small and midsize business advertisers. The ceiling for affiliate monetization was not quantified.

Cognition says Devin evolved from prototype into background-agent workflow

Cognition’s founder said AI agents have improved substantially since Devin launched about two years ago as a viral prototype with no customers. About a year later, the product worked for specific end-to-end use cases, while background-agent use was still uncommon; today, the founder said, it is becoming more commonplace, especially in their circles, as capabilities improve.

He argued that orchestration is not a “magic words” prompt trick that suddenly makes a model 15% smarter. Its practical value comes from connecting agents to the context, tools and systems required for real work.

Examples include agents that test their own code, navigate websites, inspect Datadog logs, work through large codebases and access secure company systems. The founder also pointed to combining the strengths of different models as another part of effective orchestration.

AI systems may increasingly route tasks across models instead of defaulting to the largest one

Model selection is described as a multidimensional optimization problem: systems must weigh price, speed, context knowledge, working style and specialization, not only a model’s raw logical ability. Combining models and routing each use case to the appropriate option can outperform relying on any single model.

The assessment is that this optimization will persist and may become more important as models improve. For many practical tasks, intelligence is no longer the main bottleneck, making cheaper and faster models more viable; newer, stronger models also do not necessarily produce an immediate change in ChatGPT retention metrics.

Cognition: Enterprises Are Shifting Coding-Agent Adoption Toward Measurable Use Cases

Cognition says large organizations need more time than small startups to adopt coding agents. The process includes onboarding and education, system configuration, reviews, and security guardrails.

The company says enterprise customers are moving away from measuring token consumption and are instead prioritizing specific use cases and concrete business results. These may include improving existing applications or building new capabilities, rather than only carrying out backend migrations.

Cognition argues that AI’s value should be tracked case by case, including productivity and outcomes, to distinguish effective applications from ineffective ones. It also says software quality is already improving, while further progress depends on practical implementation and distribution.

AI-Agent Economy Could Drive Demand for New Financial Infrastructure

Market commentary points to a growing model-router market: Ramp has a router, Stripe bought OpenRouter, and other players are seeking a role in the token flow.

The view is that routing is only one part of a broader market. Over the next five to 10 years, internet spending could include substantial amounts for tokens, models and agents, while payments, spend management and settlement would need to be redesigned for an agent-driven economy. Which companies will become central remains unclear, but the space could support many new products.

Cybersecurity accounts for about 10% of Devin usage, Cognition says

Cybersecurity is a small but meaningful and rapidly growing part of Cognition’s business, with the company estimating that security currently represents about 10% of Devin sessions and ACUs spent.

Cognition’s security products are about two months old. The company expects the share to grow quickly and predicts that real cyber threats will emerge as small hacker teams adopt increasingly capable AI models faster than large organizations can adapt.

Anthropic reportedly walks away from talks to acquire Descartes

Anthropic has reportedly walked away from talks to acquire Descartes, according to an exclusive Bloomberg report cited in the available evidence.

Relocation was mentioned as a possible sticking point, including whether Descartes would remain in Tel Aviv rather than move to the United States. The definitive reason the talks ended was not established.

Descartes founder Dean was described as talented, while the company’s live demonstrations of its AI image and video models were praised as unusually impressive. Descartes is expected to continue developing the technology.

Astra users apply AI to 3D creation and manufacturable designs

John Brockman said the community response to Astra had been strong and creative, describing the model as reaching “a new threshold of computer use.” Users were applying it to 3D creations and to mapping physical locations into 3D models.

He also cited efforts to use Astra for designing physical parts. In one example, someone designed a mechanism to catch hair in a shower drain and was able to manufacture it.

Brockman said everyday applications can be as important as grand scientific challenges, and described Astra’s goal as empowering individuals to solve more problems in daily life.

OpenAI is moving from chat toward proactive, agentic AI tools

OpenAI executive Greg Brockman said the company is seeing a shift from pure chat use cases toward agentic ones, while traditional chat remains widely used. He said OpenAI is working to unify these modes into systems that can act proactively and persistently according to users’ goals, with safety as a core commitment. The timing and exact form of that unification were not specified.

Brockman said ChatGPT has more than one billion weekly users, while also citing approximately 300 million weekly health queries and estimating that another one billion to 1.5 billion people have tried ChatGPT before. These figures were presented as claims without supporting detail.

The discussion identified product discovery as a major challenge: users may not realize how much capabilities have changed. OpenAI’s proposed direction includes proactive onboarding and suggestions, while specific demonstrations—such as turning an AI-generated candle-holder image into a physical object—were presented as a way to show concrete use cases and help users discover additional workflows.

Image-model advances could unlock professional design and knowledge-work markets

Recent improvements in image models are described as enabling more practical workflows. After Images 2 launched, one user spent hours designing furniture through hundreds of prompts; image models can also turn photographs of physical spaces into multiple visual variations that may guide real-world projects.

The analysis argues that professional and knowledge-work applications require a quality threshold, including precise editing, speed, creativity, diverse results and effective back-and-forth interaction. Potential uses include slide and website creation. Image generation, voice and coding are presented as capabilities that OpenAI aims to combine into one package, with further improvements potentially unlocking applications users would not have anticipated.

ChatGPT Health envisioned as a connected platform for consumers, clinicians and hospitals

OpenAI describes ChatGPT Health as having three pillars: consumer health queries, a clinician-focused service intended to provide direct citations to medical literature, and enterprise deployments sold directly to hospitals. The enterprise effort includes an integration within Epic.

The vision is for these pillars to work together, enabling information sharing across providers and potentially reducing the burden on patients to repeatedly explain their medical history. With users’ permission to entrust the system with their data, the platform could also help identify people eligible for clinical trials.

OpenAI says it is already seeing anecdotes of ChatGPT helping people double-check medical advice or identify serious issues. The discussion presents memory as a potential way to detect connections across symptoms, but does not specify safeguards, consent mechanisms or clinical limitations for the proposed platform.

OpenAI links Operator’s iterative deployment to Astra’s computer-use progress

OpenAI describes Operator as an early cloud-based computer-use system that was below the threshold for broad usefulness: it was slow, not fully accurate and painful to use. The company says real-world deployment and sustained effort helped its team address a long list of problems.

The company characterizes Astra as the result of advances in the model, computer use and voice converging. OpenAI says Astra can help across tasks people perform with computers, while acknowledging that more applications remain to be discovered.

The discussion also characterizes a shift from broad, large-company-style experimentation toward focused, startup-style execution, with the team working in the same direction. Computer-use agents operating computers as humans do are presented as a potential foundation for applications across nearly anything people can do with computers.

Forus announces $150 million Series C at $3 billion valuation

Forus announced a $150 million Series C at a $3 billion valuation, four months after announcing a $1 billion valuation.

The company says it now supports millions of people across all 50 U.S. states, is used by doctors treating patients in 85% of U.S. residential ZIP codes, and works with nine of the 15 largest global biopharma companies.

Forus pitches a nationwide network spanning the medicine lifecycle

Forus says its network is designed to support biopharma companies beyond AI-based molecule discovery, covering clinical-trial site selection and patient recruitment, launch planning, physician targeting, coverage, distribution, and post-market monitoring. The company says turning a molecule into an approved medicine can still take more than a decade and billions of dollars.

Forus says its platform is gaining nationwide visibility into medicine performance and the clinical and practice behavior of patients and physicians. As an example, it claims that 40% of people who had ever taken a new autoimmune-disease drug launched in March came through the platform.

The company characterizes this level of visibility as unprecedented and says it could make drug development and commercialization faster, more efficient, and more predictable. Its stated goal is to increase the number of medicines reaching the market each year by an order of magnitude.

Pharma leaders look to AI while the industry argues it is under-credited

Pharmaceutical leaders are described as increasingly interested in AI to reinvent their businesses and use tools that could accelerate molecule research. The same assessment says the sector’s role as an inventor and creator of medicines is poorly understood, leaving companies unfairly villainized and in need of reclaiming the story.

GLP-1 drugs are cited as a major development in obesity treatment whose creators receive too little credit. The account also claims that bariatric surgeons have been applying for technology jobs and leaving medicine because GLP-1s have effectively eliminated bariatric surgery as a meaningful specialty in the country.

Healthcare fragmentation creates an opportunity for agents

An assessment based on experience at Oscar Health identifies discontinuity and heterogeneity as central features of healthcare. There is no standard case: doctors’ offices use different systems and processes, insurers apply different rules, drugs have distinct clinical, coverage and distribution requirements, and patients have differing financial, medical and insurance circumstances.

The assessment argues that this fragmentation makes healthcare particularly suitable for agents. Technology could significantly accelerate and improve the efficiency of these processes; if one company gains visibility into and control over the full workflow, it could also create new market power and change activities historically controlled by insurers and hospitals.

Split says its bill-payment product gives consumers up to 30 days of float

Split says its SplitPay product lets consumers shift payments for housing, autos, insurance and student loans to better match their income cycles. The company says the current float is 30 days and its goal is to reach 90 days; one million people are using the product, according to Split.

The company says it built an in-house cash-flow underwriting foundation model trained on internal data and uses it instead of conventional FICO-style models. SplitPay can be used for any bill that accepts ACH, with payment timing managed dynamically.

Split says its run rate grew from $1 million in the first month to nearly $80 million. It launched with a venture-debt facility, later closed Series A and Series B rounds led by Khosla Ventures, and uses a layered set of credit facilities.

Antioch combines classical simulation with learned models for physical AI

Antioch uses a hybrid strategy to train physical autonomous systems: classical, high-fidelity simulation is applied where it works, while learned models fill gaps between simulation and reality. The company says both classical simulation and end-to-end world-model approaches currently face a substantial sim-to-real gap; classical systems can require continual additions for factors such as wind, while world models lack the large volumes of physical-world data needed for training.

Antioch believes world models will ultimately be the right approach for physical AI and that the field is moving in that direction, but says the technology is not yet fully ready for broad use. Its current approach is intended to help real companies transition toward world-model-based systems over time, alongside the company’s stated goal of enabling recursive self-improvement for physical autonomous systems. Antioch’s co-founding team met at Stanford, and one co-founder previously worked on Tesla’s Autopilot team.

Robotics timelines hinge on building real-world data flywheels

The analysis rejected a 2050 timeline for autonomous driving, pointing to Tesla, Waymo and Wave deployments that were described as working extremely well. Autonomous-driving companies benefit from a data-accumulation flywheel, including information from cases where systems do not work well.

For newer robotics and industrial-automation categories, deployments remain nascent, creating a chicken-and-egg problem. The timeline for progress is expected to depend largely on how quickly companies can collect data at sufficient scale and in the right categories.

Antioch’s proposed hybrid-simulation approach would use limited real-world data to improve simulations, then use those simulations to train better physical robots. The goal is to improve sample efficiency, accelerate deployment and create a self-reinforcing data loop.

Antioch announces Amazon Ring partnership as it targets tens of thousands of companies

Antioch announced a partnership with Amazon’s Ring team, which develops smart-security devices. The company said its present-day addressable market includes tens of thousands of companies, extending beyond humanoid robotics, autonomous vehicles and industrial automation.

Antioch’s target market is systems combining hardware, software and machine learning where real-world testing is difficult and expensive. Potential applications include robotic vacuums, camera drones and other robotic components, even when the products are not humanoid robots.

Rohan Builds an AI-Enabled Wearable for More Measurable Back-Pain Care

Rohan, who experienced back pain from age 13 for roughly a decade, built a wearable system intended to collect body data and provide more measurable physical-therapy feedback. He criticized care that lacks a clear “measure, care, measure again” process.

After moving to his childhood bedroom, Rohan built an MVP, found an early investor through outreach prompted by a Twitter post about hardware angel checks, and manufactured the product in Shenzhen. He shared the MVP on Reddit, where he received strong interest and met an early customer.

Rohan said he largely built the hardware and code himself. People with back pain and other conditions have asked how the system could be adapted for them, but its broader usefulness remains to be established.

Rohan Plans Wearable Patches for AI-Assisted Physical Therapy

Rohan plans to sell a consumer device for people with back pain that would collect data to help them assess whether different physical-therapy approaches are improving their condition. The existing hardware already produces a data feed that was connected to a 3D model for a recent demonstration.

Rohan said effective AI health decisions depend on collecting enough information about a person’s body. He is exploring patches that attach to the body and said the work has already included the wrists and fingers.

He said future sensing could extend to the head and broader body coverage, potentially collecting the biomarkers needed for day-to-day health decisions. The specific biomarkers, clinical validation and final product scope have not been established.

Back-pain wearable founder tests acquisition channels as waitlist grows

Rohan is testing multiple customer-acquisition channels for a back-pain wearable, including agency work, paid acquisition and Twitter. He said the product’s current viral moment came in part from testing Twitter, while the most effective channel has not yet been determined.

Prospective customers were joining a waitlist at yourbackhurts.com. Infomercials, podcast sponsorships and live selling were mentioned as possible direct-response approaches for reaching people with back pain.

Back-pain wearable targets January shipment and first-customer waitlist

Rohan said the company is aiming to begin shipping its back-pain wearable at the start of January. The timing remains a target, and no final price was specified.

He said extensive engineering work had reduced unit economics so the product could be offered more affordably, and said the company was trying to charge as little as possible while keeping it viable. The waitlist for first customers was to open through yourbackhurts.com.

ChatGPT Images 2.5 shows high-fidelity output but struggles with dense Waldo scenes

Early Waldo Bench results for ChatGPT Images 2.5 indicate generally high-fidelity image generation. Waldo was found in one image, while some text remained highly detailed when enlarged; however, faces became garbled under strong zoom.

The beach scene was judged too easy and not dense enough for a full Waldo test—about one-quarter of a “real Waldo” scene. The assessment described the result as possibly state of the art, but not “super intelligence” for generating Waldo scenes.

The analysis suggested that modern AI may eventually solve Waldo Bench, potentially by pairing the system with Astra, tiling, and adding more reasoning. This was an informal assessment without quantitative metrics.

Astra Converts Stop-Motion Animation Into More Consistent Video

A stop-motion-style animation made with a new image model was converted into video using Astra.

The result was described as much more consistent, with consistency highlighted as one of the major improvements. No quantitative measure of the improvement was provided.

Privacy ·