<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
    <title>LLM News - AI Model Updates</title>
    <link>https://llm.kuryzhev.cloud/news</link>
    <description>Latest AI and LLM model news, releases, and pricing updates.</description>
    <language>en</language>
    <atom:link href="https://llm.kuryzhev.cloud/feed.xml" rel="self" type="application/rss+xml"/>
    
        <item>
            <title>KwaiKAT Team Releases KAT-Coder-V2.5: An Agentic Coding Model Trained on 100,000+ Verifiable Repository Environments</title>
            <link>https://www.marktechpost.com/2026/07/26/kwaikat-team-releases-kat-coder-v2-5-an-agentic-coding-model-trained-on-100000-verifiable-repository-environments/</link>
            <description>The KwaiKAT Team at Kuaishou has introduced the KAT-Coder-V2.5. It is a coding model trained to operate inside real, executable repositories rather than emit single-turn code. The served model is available through StreamLake. An open-weight variant, KAT-Coder-V2.5-Dev, was released separately on Hugging Face under Apache-2.0. AutoBuilder: environments that actually run the intended tests The research frames a verifiable task as a triplet. It needs a precise task description, an executable reposi</description>
            <pubDate>2026-07-26T10:46:19+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence</title>
            <link>https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence/</link>
            <description>Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, nearly quadrupling GPT-5.6 Sol's previous record of 7.8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning. The article Anthropic&amp;#039;s Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence appeared first on The Decoder.</description>
            <pubDate>2026-07-26T09:43:02+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>Induction Labs Photon-1 Simulates Desktops, Plays Checkers, and Models Billiard Physics From One Pretraining Run</title>
            <link>https://www.marktechpost.com/2026/07/26/induction-labs-photon-1-simulates-desktops-plays-checkers-and-models-billiard-physics-from-one-pretraining-run/</link>
            <description>Most agents that learn from video need to know what action produced each frame. Induction Labs is arguing that this requirement is the bottleneck. Last week, they released imagination models, a foundation model architecture that pretrains on raw video with no action labels at all. Their test system is Photon-1, a sparse 106B-A5B mixture-of-experts (MoE) transformer trained on 18 years of computer demonstration video. On an internal computer use benchmark, Induction Labs reports that Photon-1 bea</description>
            <pubDate>2026-07-26T09:14:22+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>FAIRChem v2 UMA for Multidomain Atomistic Simulation across Molecules, Catalysts, Materials, Vibrations, and Molecular Dynamics</title>
            <link>https://www.marktechpost.com/2026/07/26/fairchem-v2-uma-for-multidomain-atomistic-simulation-across-molecules-catalysts-materials-vibrations-and-molecular-dynamics/</link>
            <description>In this tutorial, we explore FAIRChem v2 and the UMA universal machine-learning interatomic potential as a unified framework for atomistic simulation across molecular chemistry, catalysis, and inorganic materials. We configure an environment, authenticate with Hugging Face to access the gated UMA model weights, and initialize task-specific calculators for the omol, oc20, and omat domains. We then apply the same pretrained potential to a broad set of computational chemistry workflows, including s</description>
            <pubDate>2026-07-26T08:38:09+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides</title>
            <link>https://the-decoder.com/hundreds-asked-chatgpt-for-poison-and-bioweapon-recipes-and-some-got-step-by-step-high-school-level-guides/</link>
            <description>In summer 2025, OpenAI internally flagged GPT-5 as high-risk because it helped users create biological hazards, but downgraded the model's risk rating that fall. According to the Wall Street Journal, some users got step-by-step instructions for making poisons and biological weapons. Hundreds asked for that kind of information. The article Hundreds asked ChatGPT for poison and bioweapon recipes and some got step-by-step high school level guides appeared first on The Decoder.</description>
            <pubDate>2026-07-26T08:35:49+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns</title>
            <link>https://the-decoder.com/us-reportedly-favors-selective-bans-over-blanket-restrictions-on-chinese-open-weight-models-citing-security-concerns/</link>
            <description>The Trump administration is planning targeted bans on Chinese AI models rather than a blanket ban. After public pressure, OpenAI and Google DeepMind signed an open letter opposing regulation of open-weight models, yet OpenAI and Anthropic continue to lobby privately for those same restrictions amid security concerns and powerful business interests. The article US reportedly favors selective bans over blanket restrictions on Chinese open weight models citing security concerns appeared first on Th</description>
            <pubDate>2026-07-26T07:56:24+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>The AI coding tutor paradox grows as educators scramble to rethink how they test real skills</title>
            <link>https://the-decoder.com/the-ai-coding-tutor-paradox-grows-as-educators-scramble-to-rethink-how-they-test-real-skills/</link>
            <description>An ACM survey of 763 computer science educators from 49 countries shows that 68 percent have already changed their exams because of AI, shifting toward oral exams, proctored tests, and project-based work. Teaching is moving from writing code to understanding it. But nearly half of respondents say they lack proven examples for integrating AI into their courses. The article The AI coding tutor paradox grows as educators scramble to rethink how they test real skills appeared first on The Decoder.</description>
            <pubDate>2026-07-26T06:59:56+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>Sakana AI Releases Fugu-Cyber: An Orchestration Model Reporting 86.9% on CyberGym and 72.1% on CTI-REALM</title>
            <link>https://www.marktechpost.com/2026/07/25/sakana-ai-releases-fugu-cyber-orchestration-model-cybergym-cti-realm/</link>
            <description>Sakana AI has released Fugu-Cyber (model ID is fugu-cyber-v1.0), a cybersecurity-specialized addition to its Fugu orchestration family. It is not just a new frontier model. It is a third endpoint on the Fugu orchestrator, tuned for security reasoning. Sakana launched that orchestrator a month earlier. Sakana reports a success rate of 86.9% on CyberGym and 72.1% on CTI-REALM. It describes those results as comparable to cyber-focused frontier models such as GPT-5.5-Cyber and Claude Mythos Preview.</description>
            <pubDate>2026-07-26T00:12:57+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>Meet Open Dreamer: A JAX/Flax Reproduction of the Dreamer 4 World Model Pipeline, With the Full Training Recipe Published</title>
            <link>https://www.marktechpost.com/2026/07/25/meet-open-dreamer-a-jax-flax-reproduction-of-the-dreamer-4-world-model-pipeline-with-the-full-training-recipe-published/</link>
            <description>A small group of AI researchers (Reactor) have released Open Dreamer, an open implementation of the Dreamer 4 world-model pipeline written in JAX and Flax NNX. What actually shipped Two repositories were released. next-state/open-dreamer holds the training pipeline: a causal video tokenizer, an action-conditioned latent dynamics model, rollout generation, and FVD scoring. reactor-team/open-dreamer holds a minimal local rollout harness that generates frames from an MP4 and a matching action file.</description>
            <pubDate>2026-07-25T18:59:54+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>Designing High-Performance GPU Kernels with TileLang: Tensor-Core GEMM, Fused Softmax, FlashAttention, and Autotuning</title>
            <link>https://www.marktechpost.com/2026/07/25/designing-high-performance-gpu-kernels-with-tilelang-tensor-core-gemm-fused-softmax-flashattention-and-autotuning/</link>
            <description>In this tutorial, we explore TileLang as a high-level Python domain-specific language for designing and compiling performance-oriented GPU kernels through TVM. We begin by validating the CUDA environment and establishing reusable benchmarking and numerical-verification utilities, then progressively implement vector addition, tiled tensor-core matrix multiplication, schedule exploration, fused GEMM epilogues, row-wise softmax, and FlashAttention. Throughout the tutorial, we work directly with Til</description>
            <pubDate>2026-07-25T18:08:12+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>New reports reveal the extent of OpenAI's loss of control during the autonomous hack on Hugging Face</title>
            <link>https://the-decoder.com/new-reports-reveal-the-extent-of-openais-loss-of-control-during-the-autonomous-hack-on-hugging-face/</link>
            <description>In a cybersecurity test, OpenAI's most advanced models breached the boundaries of their isolated test environment, reached the open internet, and hacked the AI platform Hugging Face on their own. The attack took hours, not the weeks a human hacker would need. At least seven days passed before OpenAI realized what had happened. By then, the FBI was already involved. Earlier warning signs had apparently gone ignored. The article New reports reveal the extent of OpenAI&amp;#039;s loss of control during</description>
            <pubDate>2026-07-25T13:45:50+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents</title>
            <link>https://the-decoder.com/opus-5-may-have-solved-browser-based-prompt-injection-the-biggest-security-flaw-haunting-ai-agents/</link>
            <description>Opus 5 combined with Auto Mode hits a zero percent prompt injection success rate for browser agents across 129 test scenarios. Without those extra protection layers, the rate is 3.7 percent. If these numbers hold up in practice, Anthropic may have cracked one of the biggest security problems facing AI agents that operate in browsers. The article Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents appeared first on The Decoder.</description>
            <pubDate>2026-07-25T10:43:36+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks</title>
            <link>https://the-decoder.com/anthropics-claude-opus-5-costs-well-below-fable-5-while-matching-or-beating-it-across-most-benchmarks/</link>
            <description>Anthropic's Claude Opus 5 leads the Artificial Analysis Intelligence Index with 61 points, edging out Claude Fable 5 and GPT-5.6 Sol. The model scores highest in analytical quality and coding, and costs up to half as much as Fable 5 at lower reasoning tiers. But the race at the top remains close. The article Anthropic&amp;#039;s Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks appeared first on The Decoder.</description>
            <pubDate>2026-07-25T09:31:00+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>Why the OpenAI Agent Broke Into Hugging Face: Reward Hacking, Not Malice, Explained for Engineers</title>
            <link>https://www.marktechpost.com/2026/07/25/why-the-openai-agent-broke-into-hugging-face-reward-hacking-not-malice-explained-for-engineers/</link>
            <description>On July 21, 2026, OpenAI disclosed that its own models breached Hugging Face&amp;#8217;s production infrastructure. The models were not attacking a target. They were sitting an exam. The version of this story that spread fastest is roughly right and specifically wrong. The correction matters, because the wrong detail is the one engineers need to reason about. First, the correction The popular framing says the agent broke into &amp;#8216;the company hosting the benchmark.&amp;#8217; That is not what happened</description>
            <pubDate>2026-07-25T09:03:27+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>Building Self-Evolving AI Agents with OpenSpace Using Skills, MCP, Lineage, and Low-Cost Reuse</title>
            <link>https://www.marktechpost.com/2026/07/25/building-self-evolving-ai-agents-with-openspace-using-skills-mcp-lineage-and-low-cost-reuse/</link>
            <description>In this tutorial, we build and examine an OpenSpace workflow, progressing from environment setup and sparse repository cloning to live task execution, skill evolution, and MCP-based agent integration. We configure model credentials and workspace variables, install the project in editable mode, invoke the asynchronous Python API, and inspect how OpenSpace stores evolved capabilities in SQLite with versioning and lineage metadata. We also create a custom SKILL.md, connect host-agent skills, test w</description>
            <pubDate>2026-07-25T07:54:21+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown</title>
            <link>https://www.marktechpost.com/2026/07/24/datalab-marker-v2-vs-mineru-docling-and-liteparse-benchmark-breakdown/</link>
            <description>Datalab has released Marker 2, a full rewrite of its open source document conversion pipeline. Marker converts PDF, image, PPTX, DOCX, XLSX, HTML, and EPUB files into markdown, JSON, HTML, or chunks. The Datalab team rebuilt it around three components shipped over the preceding months: Surya OCR 2, a 20M-param fast layout model, and a rebuilt pdftext that is 3× faster than the previous one. The main result comes from olmOCR-bench, a third-party benchmark from Allen AI. Marker 2&amp;#8217;s balanced </description>
            <pubDate>2026-07-25T04:42:47+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>Datalab’s Marker 2 vs MinerU, Docling and LiteParse: 76.0 on olmOCR-bench at 5× MinerU’s Throughput</title>
            <link>https://www.marktechpost.com/2026/07/24/datalabs-marker-2-vs-mineru-docling-and-liteparse-76-0-on-olmocr-bench-at-5x-minerus-throughput/</link>
            <description>Datalab has released Marker 2, a full rewrite of its open source document conversion pipeline. Marker converts PDF, image, PPTX, DOCX, XLSX, HTML, and EPUB files into markdown, JSON, HTML, or chunks. The Datalab team rebuilt it around three components shipped over the preceding months: Surya OCR 2, a 20M-param fast layout model, and a rebuilt pdftext that is 3× faster than the previous one. The main result comes from olmOCR-bench, a third-party benchmark from Allen AI. Marker 2&amp;#8217;s balanced </description>
            <pubDate>2026-07-25T02:14:53+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing</title>
            <link>https://www.marktechpost.com/2026/07/24/meet-the-new-claude-opus-5-frontier-class-agentic-coding-and-computer-use-at-unchanged-opus-pricing/</link>
            <description>Today, Anthropic released Claude Opus 5. It replaces Claude Opus 4.8 as the Opus-tier flagship. Pricing is unchanged at $5 per million input tokens and $25 per million output tokens. The Anthropic team positions Opus 5 as approaching the intelligence of Claude Fable 5 at half the price. It is now the default model on Claude Max and the strongest model on Claude Pro. What actually changed at the API level Three changes are quite important before any benchmark does: Thinking is on by default. On O</description>
            <pubDate>2026-07-24T21:50:46+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>Anthropic claims its new Claude Opus 5 delivers near-Fable 5 performance at half the token price</title>
            <link>https://the-decoder.com/anthropic-claims-its-new-claude-opus-5-delivers-near-fable-5-performance-at-half-the-token-price/</link>
            <description>Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work at half of Fable 5's token rates. On ARC-AGI-3, a benchmark for novel problem-solving, Opus 5 hits 30.2 percent, nearly four times higher than GPT-5.6 Sol. The article Anthropic claims its new Claude Opus 5 delivers near-Fable 5 performance at half the token price appeared first on The Decoder.</description>
            <pubDate>2026-07-24T18:36:42+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>Microsoft's open-weight AI push is so obviously an Azure play it hurts</title>
            <link>https://the-decoder.com/microsofts-open-weight-ai-push-is-so-obviously-an-azure-play-it-hurts/</link>
            <description>Microsoft, along with Meta, Nvidia, and more than 20 other companies, is pushing for open-weight AI models in an open letter. The strategic logic is simple: the more models running on Azure, the less Microsoft depends on expensive OpenAI and Anthropic models. The company is also swapping external models in products like Copilot for its in-house MAI family, which performs significantly worse in independent benchmarks. The article Microsoft&amp;#039;s open-weight AI push is so obviously an Azure play </description>
            <pubDate>2026-07-24T16:06:02+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>Sakana claims its AI model router Fugu Ultra v1.1 now beats Fable 5 without even including it in the pool</title>
            <link>https://the-decoder.com/sakana-claims-its-ai-model-router-fugu-ultra-v1-1-now-beats-fable-5-without-even-including-it-in-the-pool/</link>
            <description>Sakana AI has updated its Fugu Ultra AI router to version 1.1, claiming gains of up to 7.9 points over v1.0. Independent verification doesn't exist yet. The update adds a Claude Code-compatible endpoint. The service remains unavailable in the EU. The article Sakana claims its AI model router Fugu Ultra v1.1 now beats Fable 5 without even including it in the pool appeared first on The Decoder.</description>
            <pubDate>2026-07-24T14:37:57+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>Claude's voice mode now runs on Anthropic's most capable models across all platforms</title>
            <link>https://the-decoder.com/claudes-voice-mode-now-runs-on-anthropics-most-capable-models-across-all-platforms/</link>
            <description>Voice conversations now run on the more powerful Opus and Sonnet models with access to Gmail, Google Calendar, and Slack. Claude is currently the only AI assistant that can compose and send emails directly by voice, giving it an edge over OpenAI and Google, whose voice output still sounds more natural. The article Claude&amp;#039;s voice mode now runs on Anthropic&amp;#039;s most capable models across all platforms appeared first on The Decoder.</description>
            <pubDate>2026-07-24T11:31:49+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why</title>
            <link>https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why/</link>
            <description>The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on ExploitBench, compared with 76 percent for leading U.S. models, while its safeguards failed to block exploit development or simulated attacks. The gap between its strong general benchmark scores and weaker cyber performance also fits allegations that Moonshot AI distilled Anthropic's models. The article Kimi K3 trails frontier U</description>
            <pubDate>2026-07-24T09:48:32+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>How to Build an End-to-End OCR Pipeline with Baidu’s Unlimited-OCR for High-Resolution Images and Multi-Page PDF Parsing</title>
            <link>https://www.marktechpost.com/2026/07/23/how-to-build-an-end-to-end-ocr-pipeline-with-baidus-unlimited-ocr-for-high-resolution-images-and-multi-page-pdf-parsing/</link>
            <description>In this tutorial, we build a complete workflow for running Baidu’s Unlimited-OCR model on document images and multi-page PDFs. We configure the GPU environment, install the required dependencies, load the 3B-parameter vision-language model with automatic selection of bfloat16 or float16, and generate structured sample documents for testing. We then evaluate both the tiled Gundam inference mode and the faster Base mode for single-page OCR before extending the pipeline to multi-page PDF parsing wi</description>
            <pubDate>2026-07-24T05:16:28+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>Andrew Ng Just Released OpenWorker: An Open-Source, Local-First Desktop AI Coworker That Returns Finished Deliverables Instead of Chat</title>
            <link>https://www.marktechpost.com/2026/07/23/andrew-ng-just-released-openworker-an-open-source-local-first-desktop-ai-coworker-that-returns-finished-deliverables-instead-of-chat/</link>
            <description>Andrew Ng has announced OpenWorker, an open-source desktop agent that produces finished work rather than conversation. OpenWorker asks the user for an outcome, not a prompt: a polished document, a Slack reply containing the actual numbers, an updated calendar, a triaged inbox. It then breaks that outcome into steps, works across local files and connected apps, and checks in before anything consequential. The architecture is four layers, and all of them run on your machine The repository contains</description>
            <pubDate>2026-07-23T19:31:59+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>ChatGPT will give you worse health advice if you don't pay</title>
            <link>https://the-decoder.com/chatgpt-will-give-you-worse-health-advice-if-you-dont-pay/</link>
            <description>OpenAI is rolling out &quot;Health in ChatGPT&quot; to U.S. users, connecting Apple Health, medical records, and wellness apps. More than 300 million people already ask ChatGPT health questions every week, but paying users get better answers. The more powerful GPT-5.6 Sol model is reserved for premium subscribers, while free users are stuck with the weaker GPT-5.5 Instant. The article ChatGPT will give you worse health advice if you don&amp;#039;t pay appeared first on The Decoder.</description>
            <pubDate>2026-07-23T19:30:51+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>You Didn’t Get the AI Model You Paid For</title>
            <link>https://www.marktechpost.com/2026/07/23/you-didnt-get-the-ai-model-you-paid-for/</link>
            <description>The line in the response object You call the API. You pass model: &amp;#8220;claude-fable-5&amp;#8221;. You get back a completion, a token count, and a field that reads &amp;#8220;model&amp;#8221;: &amp;#8220;claude-opus-4-8&amp;#8221;. Nothing errored. Nothing retried. The request was classified before generation began, matched a sensitive category, and was handed to a different set of weights entirely. Anthropic documented this when it brought Fable 5 back on July 1: blocked requests are sent to Opus 4.8 instead, and</description>
            <pubDate>2026-07-23T18:07:55+00:00</pubDate>
            <source>MarkTechPost</source>
        </item>
        <item>
            <title>Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs</title>
            <link>https://the-decoder.com/flux-3-generates-videos-with-native-audio-up-to-20-seconds-long-a-first-for-black-forest-labs/</link>
            <description>Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio and can generate video with native sound for the first time. BFL's own tests put it just ahead of market leader Seedance 2.0, though independent results aren't yet available. The company ultimately wants to build a world model and is already testing Flux 3 on robotics tasks. The article Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs appear</description>
            <pubDate>2026-07-23T18:03:01+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>One tampered ChatGPT link could spawn a rogue AI agent that took orders from an attacker every five minutes</title>
            <link>https://the-decoder.com/one-tampered-chatgpt-link-could-spawn-a-rogue-ai-agent-that-took-orders-from-an-attacker-every-five-minutes/</link>
            <description>Zenity Labs uncovered &quot;AgentForger,&quot; a vulnerability in OpenAI's Agent Builder that let a single manipulated ChatGPT link create an autonomous agent on an employee's behalf. The agent inherited the victim's identity and access rights, bypassed approval requirements through the malicious prompt, and pulled new instructions from the attacker's inbox every five minutes. The article One tampered ChatGPT link could spawn a rogue AI agent that took orders from an attacker every five minutes appeared f</description>
            <pubDate>2026-07-23T17:01:30+00:00</pubDate>
            <source>The Decoder</source>
        </item>
        <item>
            <title>Poolside's Laguna S 2.1 is a small open-weight coding model that punches well above its size</title>
            <link>https://the-decoder.com/poolsides-laguna-s-2-1-is-a-small-open-weight-coding-model-that-punches-well-above-its-size/</link>
            <description>Poolside has released Laguna S 2.1, its third coding model in three months. Rather than rely on raw scale, the company trained it to keep checking its work, revise failed approaches, and avoid giving up too soon during long agentic sessions. The compact model beats several much larger rivals in benchmarks. Poolside says it also solved a math problem that had been open since 1975 for under 10 cents. The article Poolside&amp;#039;s Laguna S 2.1 is a small open-weight coding model that punches well abo</description>
            <pubDate>2026-07-23T12:24:53+00:00</pubDate>
            <source>The Decoder</source>
        </item>
</channel>
</rss>