Andrew Ng Just Released OpenWorker: An Open-Source, Local-First AI Desktop Partner That Returns End-to-End Delivery Instead of Conversation

Andrew Ng Just Released OpenWorker: An Open-Source, Local-First AI Desktop Partner That Returns End-to-End Delivery Instead of Conversation

Andrew Ng announced OpenWorker, an open source desktop agent that generates a finished task rather than a conversation. OpenWorker prompts the user i the resultnot informed: polished document, Slack response containing real numbers, updated calendar, tripled inbox. It then breaks that output down into steps, runs through all local files … Read more

Anthropic Releases Code Claude Security Plugin in Beta: Multi-Vulnerability Agent Scanner That Works on Your Terminal

Anthropic Releases Code Claude Security Plugin in Beta: Multi-Vulnerability Agent Scanner That Works on Your Terminal

Anthropic released the Claude Security plugin for Claude Code in beta. The plugin runs multi-agent vulnerability scans in the repository within an existing Claude code session, and converts selected findings into patch files that you update and apply. Anthropic emphasized the tool’s interoperability in a variety of ways after the announcement, highlighting its ability to … Read more

Cursor Releases Cursor Router: An Application-Level Classifier Delivers Frontier Code Quality at 30–50% Lower Cost

Cursor Releases Cursor Router: An Application-Level Classifier Delivers Frontier Code Quality at 30–50% Lower Cost

Cursor has made Cursor Router widely available for Teams and Enterprise applications. The system is a system that evaluates each request before the model runs, and then sends it to the model best suited for that particular task. The cursor team reports quality performance bordering on 60% savings in online A/B testing, and 30–50% savings … Read more

EdgeBench Research Grade Analytics: AI rating agent, Leaderboard Analytics, Ranking rules, and evaluation metrics

EdgeBench Research Grade Analytics: AI rating agent, Leaderboard Analytics, Ranking rules, and evaluation metrics

In this lesson, we explore EdgeBench as a practical benchmark for evaluating advanced AI agents across various task categories, runtime environments, and interaction time budgets. We begin by downloading a dataset summary from Hugging Face, dissecting the extracted task specification, and examining the benchmark taxonomy, performance settings, online requirements, judgment logic, and score metadata. We … Read more

Cisco Foundation AI Releases Antares: 350M and 1B Open Weight Models Reveal Known Vulnerabilities Inside Real Codes

Cisco Foundation AI Releases Antares: 350M and 1B Open Weight Models Reveal Known Vulnerabilities Inside Real Codes

Cisco Foundation AI released Antaresa family of security small language models (SLMs) designed for one small security task. The task is to localize vulnerability. Given a description of the vulnerability and the cache, find the files that contain the bug. Two open-air models are also available now at Hugging Face, Antares-350M and Antares-1B. Both are … Read more

Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model That Punches Above Its Weight Class on SWE-Bench Multilingual

Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model That Punches Above Its Weight Class on SWE-Bench Multilingual

Poolside has released the Laguna S 2.1, a model 118B-parameter open-weight model designed for agent coding. It is a Mixture-of-Experts (MoE) model with 8B activated parameters per token. It supports a context window of up to 1M tokens in both logic and logic modes. The weights are in Hugging Face … Read more

Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A Cheaper, More High-Performance Flash Tier Built for Agentic Workloads

Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A Cheaper, More High-Performance Flash Tier Built for Agentic Workloads

Developers building production agents need high token efficiency, low latency, and high reliability. Today, Google released three new Gemini models. The list is Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. All three reside in the Flash category, which Google favors for speed, cost, and high-volume agency … Read more

Validating Estimates Using Distributed LLM with NVIDIA srt-slurm, SLURM Recipes, Parameter Sweeps, and Pareto Analysis

Validating Estimates Using Distributed LLM with NVIDIA srt-slurm, SLURM Recipes, Parameter Sweeps, and Pareto Analysis

In this lesson, we explore srt-slurm for NVIDIA framework and learn how we use srtctl to convert our declarative YAML configuration into a reproducible workflow for the SLURM benchmark for distributed LLM deployment. We set up the project in Google Colab, explore its internal architecture, define the cluster configuration, built-in and custom recipes, and model … Read more

Meta Open-Sources Astryx: An Agent-Ready React System with 150+ Accessible Components, Seven Themes, and a CLI.

Meta Open-Sources Astryx: An Agent-Ready React System with 150+ Accessible Components, Seven Themes, and a CLI.

Meta released Astryx, an open source design system that is fully customizable and built for use by both humans and the AI ​​agents working around them. Now available in Beta. Astryx is not a new experiment. It grew within Meta over the past eight years, when the company says it became its most widely used … Read more