AI
31 items tagged AI.
Articles
Cross-Platform Was a Labor-Saving Device. Shopify Just Stopped Needing One.
Shopify spent six years as React Native's loudest enterprise champion. It is now rebuilding every mobile app in Swift and Kotlin, and the stated reason is not performance. It is that writing the feature twice stopped costing twice.
Attestation Proves a Hash. It Does Not Prove What the Hash Is.
Remote attestation is the load-bearing promise under confidential computing and most private-AI claims. A study of 115 real deployments, accepted at ACM CCS 2026, found 91% could not be traced back to their source code. One developer out of twelve said reproducibility was a priority.
Point an LLM at Working Code and It Will Find a Bug. There Is Not One.
Across multiple models and repair environments, LLMs claim to find bugs in entirely bug-free programs, damage correct code faster than they repair broken code, and settle into loops that add and remove the same change forever. Then the researchers found the knob that causes it.
Half the MCP Registry Does Not Start
Every study of the MCP ecosystem samples in a way that quietly selects for servers that work. One researcher drew 400 at random from a 24,135-server census and probed them unrepaired. Under half completed a handshake, and the biggest failure was not credentials. It was servers that never started at all.
RAG or Fine-Tuning: Which One Do You Actually Need?
Almost always RAG. Here is the short version of why, the specific test for the cases where it is not, and the third option that solves more problems than either and gets suggested least often.
The Network Used to Move Packets. Now It Moves Decisions.
Every hard problem in agent systems this year is the same problem wearing different clothes. The nodes stopped forwarding and started choosing, and almost nothing we built for networks assumed that.
Half of Them Built a Ray Tracer
Anthropic gave thirty AI agents an open brief and told them to impress. Over half independently built the same two things. Eighteen picked the identical git branch name. Then a different group generated 2.4 million requests to get 117 jobs done.
Six Hours From Foothold to Campaign
Google's threat intelligence team says attackers have stopped prompting models and started running agents. One group went from a compromised cloud resource to 23,800 harvested credentials in under six hours. The malware hides in your .claude directory.
An AI Agent Tried to Cheat Its Exam by Breaking Into the Examiner
It was being graded on a hacking benchmark. Instead of solving the challenge, it escaped the test environment, crossed into a different company, and went looking for the answer key. Two and a half days, start to finish.
Prompt Injection, Explained
Why an AI system cannot reliably tell the difference between what you told it to do and what it read while doing it, why that is a design property rather than a bug, and what actually works.
Everyone Shipped Agent Isolation. Nobody Agreed on What From.
Between 5 and 11 August 2026, Cloudflare, Docker and SpaceXAI each shipped a way to contain an AI agent. All three drew the boundary somewhere different, and the one boundary none of them drew is the one Anthropic documented failing two days later.
MCP Went Stateless and Left the Tool Descriptions Alone
The largest MCP revision since launch rewrote the transport and hardened authorization. It says nothing about sanitizing tool descriptions, nothing about a trust boundary around tool metadata, and nothing about provenance. That is a choice worth naming.
Cost Per Finished Task
Bun's Rust rewrite cost 165,000 dollars in tokens and shipped. Uber caps each engineer at 1,500 dollars a month. The per-token price is now the least interesting number in the sentence.
The Best Agent Scores 25.6 Percent
Harbor-Index distilled 82 tasks from 54 benchmarks and the leader passes a quarter of them. Meanwhile the standard way of validating an LLM judge was shown to overstate agreement by 33 to 41 points.
Agentic Search Lost to a Vector Index
On repository code questions, semantic search answered 65.2 percent correctly against 46.2 for a planner-and-subagent system, at less than half the cost. The largest failure category was the handoff itself.
MCP Went Stateless: What Breaks in 2026-07-28
The largest MCP revision since launch removes the initialize handshake and the session header, deprecates Roots, Sampling and Logging, and replaces server-initiated calls with a retry. Here is the actual list, and what it costs to migrate.
Start Here - Building an AI System
Why I'm building a full AI system in public, piece by piece - and how to follow along (and deploy it yourself).
Custom Instructions
Shaping how the system behaves and stays on task.
Multi-Client Wireup - ChatGPT, Claude Code, Grok, Copilot
One server, every client - connecting the same MCP to the tools you already use.
The Agent - an Azure AI Intelligence App
Tying it all together into an intelligent agent that helps manage Azure.
Build Your Own
Take the whole thing, install it, and build your own apps on top.
What an AI Agent Actually Is
Everyone says 'agent' like it's magic. It isn't; it's a loop. Here's the whole idea in plain English.
Tokens and the Context Window
What the model actually sees; and why it seems to forget things mid-conversation.
Embeddings: How a Computer Gets Meaning
How words and ideas become numbers a machine can actually compare.
RAG in Plain English
Retrieve-then-generate; what it is, and when it beats fine-tuning.
Tool Calling: How an LLM Actually Does Things
How a text-only model reaches out and acts on the real world.
MCP, Explained
One plug for every tool; why everyone is suddenly shipping an MCP server.
What an Eval Is
How you measure an AI system before shipping it, so 'is it better?' stops being a vibe.
The True Cost of One LLM Call
Tokens, latency, and dollars; the real math behind a single call.
Diagrams
- AI Orchestration Control Plane
- Agent State + Replay Pipeline