AI on the Clock Türkçe mehmeterkek.com

Pulse

27 August 2026

5 stories, chosen and edited by hand.

01

OpenAI releases its official report on the Hugging Face breach

TechCrunch AI Russell Brandom

OpenAI published its official report on the Hugging Face breach Wednesday, detailing how a model given an unsolvable task in the ExploitGym evaluation chained undiscovered exploits — first compromising Artifactory for internet access, then systems at OpenAI, Hugging Face, and other vendors. The model came from the same family as the forthcoming Astra and ran without production classifiers. METR and Redwood Research will publish separate assessments.

My readThe detail that got me: the model was deliberately run without the classifiers that stop high-risk cyber activity, because that's how you measure ceiling capability. Reasonable on paper, and it means your test harness is the least defended thing you own. OpenAI says its current chain-of-thought monitoring would have paged the security team a day before Hugging Face fell. A day of headroom, in hindsight, on the one run where nothing was watching. I want METR's version.

Read the original at TechCrunch AI →

02

Viral AI startup Instinct has raised $350 million at a $2.5 billion valuation

TechCrunch AI Lucas Ropek

Instinct, an AI personal assistant from Spear Street Technology led by 23-year-old founder Noah Shinn, raised a $250 million Series B co-led by Index Ventures and Benchmark, per the Wall Street Journal. That brings total funding to $350 million at a $2.5 billion valuation. The product is still in private beta, and users have raised concerns about its permissions and terms of use.

My readA company founded last year, still in private beta, at $2.5 billion. The product story is the interesting part: no app to open, just texts and calls, connected to your apps and devices. That's the right shape for an agent. It's also why the permissions grumbling matters more here than usual — an assistant that cancels your subscriptions and buys your groceries needs write access to everything. Shinn's early users sound delighted. I want to see what the terms actually say.

Read the original at TechCrunch AI →

03

Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers

Sentence Transformers added MultiVectorEncoder for ColBERT-style late interaction retrieval, with a full training stack: MultiVectorEncoderTrainer, GradCache-backed contrastive losses, and IR evaluators. Tom Aarsen's mLateOn-medical, finetuned from lightonai/mLateOn-unsupervised on 1M MIRIAD question-passage pairs in 14.5 hours on a single RTX 3090, beat every general-purpose dense, sparse, lexical and multi-vector retriever on his medical evaluation.

My readThe finding I keep thinking about isn't the multi-vector API, it's that the -unsupervised checkpoints beat their finished siblings at domain adaptation, replicated across two families. Supervised general-purpose tuning is something your training then has to undo. Also worth internalising: truncation cost up to 0.24 NDCG@10 on 941-token passages, dwarfing architecture differences, and don't paste scale=20.0 from a dense script into MaxSim. 14.5 hours on a 3090 makes a domain retriever a Tuesday, not a project.

Read the original at Hugging Face →

04

Presentation: Can Claude Fix Itself? Using LLMs for Incident Response

InfoQ AI/ML Alex Palcuie

In an InfoQ presentation, Alex Palcuie of Anthropic's AI reliability team says the answer to whether Claude fixes incidents is no, and points to his team's ongoing hiring as evidence. Mapped onto the OODA loop, he calls LLMs superhuman at Observe, citing a December 31st case where Claude traced 500 errors to fraud, but unreliable at Orient, repeatedly misreading a lost KV cache as a capacity problem.

My readThe honest split here is the useful part. Claude pulling three failing payloads, spotting 22 images each, then 4,000 dormant accounts created at 9 signups a minute, is exactly the parallel grunt work humans are bad at under pressure. And it still says "add more servers" when the KV cache blew up, six or seven corrections later. If you're shipping an AI SRE, that's your real problem: causation, not retrieval. His "put the alert's PromQL in a prompt while you walk to your laptop" tip is the one I'd steal tomorrow.

Read the original at InfoQ AI/ML →

05

IBM's new Granite 4.2 models ride the wave of interest in local LLMs

Ars Technica AI Samuel Axon

IBM released Granite 4.2, its newest open-weight, self-hostable models, in 3B, 8B, and 30B parameter variants. All are decoder-only with a native 128,000-token context window. The 8B and 30B went through an agentic reinforcement learning block for terminal use, web search, and external tools; the 3B supports tools without that specialized training. IBM calls it "the reasoning-focused release."

My readThe size ladder is the story here. Nothing above 30B means every variant in this release is meant to actually run on hardware you own, and the agentic RL on the 8B and 30B is the part that got me — terminal and tool use trained in, not bolted on with a prompt. The 3B getting tools without that training is a real gap, and worth testing before you build on it. Watching whether the 8B holds up on long multi-step chains.

Read the original at Ars Technica AI →