Build a WebLLM Browser Inference MCP Server for Edge-Deployed Agent Reasoning in 2026
WebLLM by mlc-ai runs 27B parameter models entirely in-browser at 45 tok/s via WebGPU. Build an MCP server that exposes browser-native LLM inference as standard MCP tools for zero-infrastructure edge agents.