·
SOURCES
TRENDING · LAST 24H
  • ·Gemini 4 Argon released with 1M token context window for cybersecurity, legal, and software engineering workflows
  • ·WhiteMatter architecture achieves comparable performance with 50% more layers while using half the KV cache
  • ·Oído 13M parameter Conformer-CTC model achieves 8.4 mean WER on ESP32-S3 microcontrollers, beating Whisper-tiny
  • ·ElevenLabs hits $22B valuation via $300M tender offer; OpenAI revenue run rate reportedly nears $70B
  • ·vLLM adds PCP pipeline parallelism support; llama.cpp adds ModernBERT reranker support via classifier_pooling
#1[HN]
·
7h ago
Google DeepMind Releases Gemini 4 Argon for Professional Workflows
72 pts · 17 comments

Gemini 4 Argon introduces a 1 million token context window optimized for complex, long-horizon reasoning. The model focuses on high-stakes professional domains including autonomous cybersecurity patching, legal drafting, and software engineering.

breakdown →
#2[DEEPMIND]
·
13h ago
SynthID Bio Proof of Concept for Protein Watermarking

SynthID Bio demonstrates a method for embedding digital watermarks into AI-designed protein sequences. The approach maintains biological function while providing a mechanism to identify synthetic origins.

breakdown →
#3[THEVERGE]
·
13h ago
Google Pilots Revenue Sharing with Publishers for AI Search

Google is testing a pilot program with approximately 100 publishers to pay for content used in AI-powered search features. The move addresses publisher concerns regarding traffic loss caused by generative AI overviews.

breakdown →
#4[HUGGINGFACE]
SoL-Refiner for Single-Step 4K Video Generation

SoL-Refiner uses a three-stage recipe involving continual training, RL, and distillation to transform low-resolution video into 4K in a single denoising step. This eliminates the second sampling bottleneck common in traditional multi-step refinement.

breakdown →
#5[TLDR]
17h ago
Lemma Open-Source Workspace Enables Multi-Session Agent Collaboration

Lemma provides an open-source environment for shared, permissioned records. The platform allows human users and AI agents to maintain collaborative state across multiple sessions.

breakdown →
#6[GH]
13h ago
vLLM Adds Support for Pipeline Parallelism with PCP
★ 0 new · 0 total

The vLLM inference runtime now supports Pipeline Parallelism (PP) using PCP in the GPU Model Runner V2. This update enhances the ability to distribute large model workloads across multiple GPUs more efficiently.

breakdown →
#7[arXiv]
·
20h ago
Mnemon Memory Agent Splits LLM Work into System 1 and 2
cs.CL, cs.AI, cs.IR

Mnemon is a long-term memory architecture that bifurcates memory management into fast System 1 judgments and slow System 2 planning. A fast decision model, Jev, performs rapid assessments of raw records, while a slower LLM manages complex search queries and response composition.

breakdown →
#8[r/LocalLLaMA]
10h ago
Oído Speech Recognition Outperforms Whisper-tiny on $5 Microcontrollers
112 upvotes · 27 comments

Oído uses a 13M parameter NVIDIA Conformer-CTC Small model to achieve 8.4 mean WER in noisy environments on an ESP32-S3. This outperforms Whisper-tiny.en on a laptop, running entirely on a microcontroller without a GPU or NPU.

breakdown →
#9[HN]
·
10h ago
Strata Semantic Layer Enables Controlled LLM Data Access
3 pts · 0 comments

Strata is a full-stack semantic layer designed to provide LLMs with controlled access to business data. It balances expressiveness with ease of use, allowing non-technical users to interface with LLMs through structured dashboards and data exports.

breakdown →
#10[OPENAI]
8h ago
OpenAI Mitigates Adversarial Model Distillation Attacks

OpenAI identified and disrupted a coordinated campaign designed to extract protected model reasoning via distillation. The company is implementing new defensive measures to harden models against these specific adversarial extraction techniques.

breakdown →
#11[TECHCRUNCH]
8h ago
ElevenLabs Valuation Reaches $22B Following $300M Tender Offer

AI voice startup ElevenLabs has doubled its valuation to $22 billion. The employee tender offer was co-led by Wellington and T. Rowe Price, signaling massive institutional interest in generative audio.

breakdown →
#12[HUGGINGFACE]
FRAC Architecture Uses Fractional Dynamics for Long-Sequence SSMs

FRAC replaces exponential decay in State Space Models with power-law long memory derived from fractional dynamics. It uses a log-spaced sum of exponential modes to enable efficient parallel training and autoregressive decoding.

breakdown →
#13[GH]
13h ago
Model Context Protocol GitHub Server Repository
★ 48 new · 90,738 total

The Model Context Protocol (MCP) servers repository provides implementations for connecting LLMs to external data sources and tools. It enables standardized communication between AI models and local or remote services.

breakdown →
#14[arXiv]
·
21h ago
Multilingual voice agent benchmarks reveal performance gaps in Korean and Mandarin
cs.CL, cs.AI, cs.SD, eess.AS

The tau-Multilingual benchmark evaluates voice agents across five languages, revealing that Korean and Mandarin performance drops by 14.7 and 8.4 task-completion points respectively compared to English. Failure modes include increased missed responses in Korean and higher interruption rates in Mandarin.

breakdown →
#15[r/OpenAI]
OpenAI revenue run rate reportedly approaches $70B
204 upvotes · 59 comments

Leaked data suggests OpenAI's revenue run rate is nearing $70 billion, driven by a doubling in enterprise sales.

breakdown →
#16[HN]
10h ago
Magnitude Inference Engine Runs Agents 2x Faster Than llama.cpp
3 pts · 0 comments

Magnitude is a self-optimizing inference engine designed specifically for agents, targeting Mac, Linux, and Windows. It claims up to 2x speed improvements over llama.cpp by optimizing execution for the specific hardware and agentic use cases.

breakdown →
#17[APPLE_ML]
Systematic Study Reveals Fluency Trade-offs in LLM Conditioning

Research identifies a significant trade-off between conditioning effectiveness and output fluency in LLMs. Methods used to inject or remove specific concepts often cause a steep decline in the linguistic quality of the generated text.

breakdown →
#18[TECHCRUNCH]
13h ago
Restate Secures $20M for AI Agent Infrastructure

Restate, a startup founded by Apache Flink veterans, has raised $20M to build durable workflow infrastructure. The company aims to compete with Temporal by focusing on the reliability requirements of AI agents.

breakdown →
#19[HUGGINGFACE]
ReImaGin Framework for Visual Reasoning via Image Generation

ReImaGin enables multimodal LLMs to use image generation models as flexible, natural-language-driven reasoning tools. Unlike rigid detection or depth modules, these generators can perform open-ended visual transformations to support chain-of-thought reasoning.

breakdown →
#20[GH]
17h ago
vLLM Fixes Mamba2 Quantized Weight Loading for Tensor Parallelism
★ 0 new · 0 total

vLLM updated its inference runtime to correctly load Mamba2 quantized in_proj weights and scales when using tensor parallelism (TP > 1). This fix ensures scaling factors are applied correctly across distributed GPU setups.

breakdown →
#21[arXiv]
22h ago
Systematic Study of ZK-Friendly Quantization for Verifiable LLM Inference
cs.AI

This research evaluates how quantization methods impact the arithmetic complexity and proving costs of Zero-Knowledge (ZK) proofs for LLM inference. It identifies the relationship between quantization schemes and the underlying finite field operations required for verifiable privacy and governance.

breakdown →
#22[r/LocalLLaMA]
13h ago
llama.cpp adds GLM-5.3-Flash support
116 upvotes · 38 comments

The llama.cpp repository includes a pull request for GLM-5.3-Flash (GLM5-Next) support. This integration allows users to run the model locally using GGUF quantization.

breakdown →
#23[HN]
17h ago
Pi.dev Integrates Model Context Protocol (MCP) into Core
5 pts · 0 comments

Pi.dev has pivoted from rejecting the Model Context Protocol (MCP) to integrating it directly into the platform core. This allows for native support of MCP-compliant tools and data sources instead of relying on extensions.

breakdown →
#24[APPLE_ML]
SCLATE Provides Unified Substrate for Continual-Learning Agent Evaluation

SCLATE is an execution substrate that uses an adapter to allow benchmarks and agents to share a single event scheduler. This enables standardized evaluation of long-horizon tasks involving session management and memory consolidation.

breakdown →
#25[TECHCRUNCH]
8h ago
OpenAI Decisions API Targets High-Speed Agent Coordination

OpenAI's Decisions API functions as a Jev clone, optimized for low-latency, low-cost intelligence. The tool aims to manage and regulate swarming agents by providing rapid decision-making capabilities.

breakdown →
9 sources · live pipeline status →↓ mac menu bar app (apple silicon)
·