Judge rules Trump administration’s blacklisting of Anthropic was illegal
The ruling knocks down the Trump-era ban, reopening the door for Anthropic to do business with the U.S. government — unless a higher court stays or reverses it. Expect agencies that enforced the blacklist to revisit their guidance and any procurements that excluded the company, with the immediate pace hinging on whether the government seeks an appeal or stay.
The thread centers on the legal mechanics of the ruling: because courts traditionally grant immense deference to the executive branch on "national security," the ban was overturned not just because the government's evidence was thin—relying on a mere four-page memorandum—but because that lack of evidence proved the ban was naked retaliation against protected speech.
From there, the discussion broadens into a bipartisan critique of "national security" as an all-purpose escape hatch to bypass congressional scrutiny. Commenters fiercely debated the history and structure of executive overreach:
- Historical precedents: Some argued that suspending norms for security is an American tradition dating back to the 1860s and 1940s. Others countered that equating peacetime political maneuvers (or the perpetual "War on Terror") to the existential threats of the Civil War and WWII is a false equivalence.
- Systemic design: The conversation fractured over whether the expanding executive branch is a fatal flaw in the presidential system—allowing the president to effectively legislate without Congress—or if the Founders' system is sound but failing because modern voters and institutions refuse to use constitutional tools to punish bad-faith leadership.
- Corporate speech: A minor but pointed observation surfaced the irony of the case: the defense against the administration hinged entirely on corporate First Amendment rights, a legal doctrine frequently criticized on Hacker News but which provided the ultimate shield here.
GLM-5.3 is now open-weight
All improvements come from post-training on the same base as GLM‑5.2, yet they report a 50% coding jump on their in‑house Z.ai Code Bench and open‑source SOTA on Terminal Bench 3.0 and Agents’ Last Exam. Concrete deltas vs 5.2: Terminal Bench 3.0 avg@3 climbs 4.6 → 28.3; CyberGym 77.2 → 84.5; ExploitGym Pass@1 (2h/6h) 29/39 → 105/130; ExploitBench 24.4 → 54.4.
Weights are downloadable and run locally via SGLang, vLLM, TokenSpeed, Transformers, KTransformers, and Unsloth; Ascend NPU inference is supported (vLLM‑Ascend, xLLM, SGLang). There’s a knob to control compute spent on reasoning: reasoning_effort={low, high, max} (default max). For chat, explicitly set clear_thinking=true.
Methodology notes: most results are with their Claude Code 2.1 harness and long contexts; some are single‑run Pass@1 with generous or model‑TPS‑rescaled time budgets (e.g., ExploitGym).
The thread is entirely consumed by a debate over the economics of buying high-end local hardware—such as upcoming 512GB Macs or multi-GPU rigs—versus relying on cloud APIs to run open-weight models.
The API pragmatists argue that cloud economies of scale have rendered local hosting financially irrational, estimating a ten-year payback period for high-end gear. They emphasize that fierce competition across platforms like OpenRouter keeps prices low and ensures legacy models aren't entirely sunsetted. One user admitted their expensive Strix Halo and dual-32GB GPU setup now sits idle, as the sheer electricity and cooling costs during a Texas summer nullified any savings compared to hitting cloud endpoints.
In the opposing camp, local hosting advocates argue that absolute privacy and infrastructural "object permanence" justify the capital expense. They dismiss cloud providers' Zero Data Retention (ZDR) policies as a "pinky promise" that still fundamentally requires transmitting unencrypted data—a risk they note is compounding as autonomous coding agents increasingly scrape messy local terminal environments and system files. For these developers, owning the metal is the only verifiable hedge against arbitrary API rate hikes, silent model alterations, or the exposure of highly profitable niche workflows.
Despite the deep skepticism toward dropping cash on next-generation hardware, older Apple Silicon remains the community's acknowledged sweet spot: M1 Max owners reported persistent satisfaction, achieving 50–60 tokens per second on ~30B parameter models. The overriding consensus leans toward holding off on major purchases, as inference software currently optimizes faster than silicon, and rumors of ultra-efficient upcoming architectures suggest the hardware baseline is about to shift.
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Across 12 AlphaEvolve construction problems (plus two case studies), agents delivered novel results on five, including a new infinite family of finite-field Kakeya sets; exact 604-point kissing configurations in 11 dimensions; new records for the discretized Kakeya needle and sign uncertainty problems; and a substantially improved lower bound for Erdős’s minimum-overlap problem, along with novel infinite families for Book Ramsey numbers.
The work runs in “the Station,” an open-world multi-agent setup where models from different families self-select research directions, run experiments, collaborate, and build a shared literature — all without a central coordinator or scripted pipeline. Importantly, the agents produced not just numerical constructions but theorems and analyses explaining why they work, improving interpretability and handoff to human mathematicians. The authors publish raw agent dialogues, proofs, and verification code to provide a transparent trace of how each discovery emerged.
The thread's core debate centered on the nature of AI creativity and whether these results finally satisfy skeptics who demand "novel" mathematical discoveries. While some users saw this as definitive proof of original work, others argued this misconstrues the mathematical community's actual critique: that AI's success currently stems from the rapid testing and recombination of a vast memory. This constitutes a valid but potentially non-exhaustive form of creativity, leaving it an open question whether the system's approach can cover all forms of mathematical intuition.
Other discussions focused on the paper's framing and system architecture:
- Anthropomorphism: The authors' description of giving agents "holidays" for open-ended thought sparked a philosophical debate. Users weighed whether using human cognitive terms distorts technical understanding or helpfully demystifies human intelligence.
- Endogenous institutions: One commenter proposed extending the environment by allowing agents to build their own reputation systems, journals, and peer-review standards rather than relying on architect-defined rewards. A counterpoint noted that the limited context windows of current models might make agents too "short-lived" to develop meaningful institutional status signals.
- Literature and Code: Multiple readers compared the multi-agent setup to the "truth mines" in Greg Egan's hard sci-fi novel Diaspora, and users surfaced the underlying open-source repository at
dualverse-ai/station.
Luanti removed from Google Play due to baseless AI copyright notice
An AI brand-protection platform working for Microsoft filed a vague DMCA that got Luanti’s Android app pulled from Google Play, without identifying any specific Minecraft assets. The notice cites US Reg. #TX 8-192-097 (Minecraft Java Edition 1.9) but doesn’t say what Luanti allegedly uses. Luanti says the app ships with no games or third‑party game assets—only a small set of engine textures and fonts with proper attribution—and it stopped bundling Minetest Game in December 2023; community content is fetched from ContentDB, where uploads are manually reviewed for licensing. They previously beat a near-identical 2023 notice from the same company, which also targeted another voxel-style indie game (Allumeria). The project argues voxel “cubes” are a genre, not proprietary, and that any valid complaints should target specific ContentDB packages, not the engine itself; they’re evaluating perceptual hashing to assist moderators but insist on human review. The incident highlights how automated takedowns, combined with platform safe-harbor workflows, can sideline open-source apps absent concrete evidence.
The dominant correction in the thread is that Luanti was not actually hit by a formal DMCA notice, but rather Google’s proprietary, parallel dispute-resolution system. Commenters explain that platforms rely on these extralegal "pseudo-DMCA" workflows to appease legacy media partners and sidestep the formal DMCA requirement to restore content after receiving a counter-notice.
The discussion highlights the asymmetric warfare baked into the current enforcement ecosystem:
- The jurisdiction trap: Even if developers can force a formal DMCA counter-claim, international creators are legally required to consent to US Federal Court jurisdiction (typically in California). The sheer cost of hiring US defense counsel to survive a motion to dismiss—especially in a visual gray area like voxel engines—is an intentional deterrent that forces targets to settle.
- The economics of automation: While some users call for Microsoft to fire the attorneys overseeing Tracer.AI, others point out that automated dragnets exist precisely to save corporate legal fees. Microsoft is effectively offloading the cost of false positives onto indie developers, a problem that would theoretically require hiring more human reviewers, not fewer.
- A proposed deterrent: One user outlined a tiered bond system where high-volume, automated copyright enforcers would have to post escalating financial bonds, which would be directly forfeited to falsely accused creators when a strike is reversed.
Ultimately, the thread views the ordeal less as a strict failure of copyright law and more as the inevitable result of closed platforms operating unregulated, private justice systems designed to minimize their own liability.
Run Qwen3.8 27B locally: real numbers from my Mac Studio
On a Mac Studio M3 Ultra, Qwen3.8-27B (Q4_K_M, ~17GB) averages ~14 tok/s via Ollama—about half qwen3.6:27b’s ~28.6 tok/s—yet reaches similar wall-clock time per answer because it uses ~1/3 the tokens. The author timed five runs per model on identical prompts (200–500 word outputs): 3.6 typically used 1,950–3,340 tokens vs 3.8’s 890–1,090, making 72s at 28.6 tok/s vs 67s at 14.2 tok/s a practical tie. Prompt processing throughput is similar (95 vs 93 tok/s), but generation speed diverges; the new hybrid attention (“qwen35”) likely hasn’t been fully optimized in Metal yet.
Both tests used Ollama’s default Q4_K_M (each ~17GB on disk) on an M3 Ultra with 256GB RAM. While generating, inference pinned all 60 GPU cores at 100% (GPU ~64W) with the CPU around 6W; system draw peaked near 291W—pure GPU work.
The 1‑bit Unsloth quant (6.7GB) was fast—~309 tok/s prompt processing and ~27.2 tok/s generation—and got factual recall right, but it wouldn’t commit: for a simple bash one‑liner it spun ~400 tokens second‑guessing itself. This matches Unsloth’s guidance: 1‑bit is not for agentic/tool‑calling; their stated floor for that is Q2_K_XL (~9.8GB). The takeaway: quantization doesn’t degrade evenly—facts survive, decisiveness dies.
Practical notes: 32GB RAM comfortably runs Q4; 16GB runs Q2. You need a very recent llama.cpp build; older ones error with unknown model architecture “qwen35.” The per‑token slowdown on 3.8 should shrink as runtimes catch up, but today its thriftier outputs already neutralize the speed gap in end‑to‑end use.
The thread immediately flags a flaw in the benchmark: Qwen 3.8 and 3.6 share the exact same architecture and parameter count, meaning the 50% generation slowdown is a software anomaly, not an inherent model trait. Commenters suspect an Ollama bug, a botched GGUF conversion, or MTP (multi-token prediction) mispredictions. To bypass the bottleneck, power users recommend serving the model via oMLX with internal MTP enabled, which can push throughput back up to 40–90 tok/s on high-end M-series chips—though they warn against enabling Dflash, which can break batching.
The rest of the discussion fractures into specific hardware and stack recommendations:
- Budget inference rigs: Users running older AMD datacenter cards (like the 32GB MI50 or dual MI25s) report hitting 30–70 tok/s using llama.cpp with Vulkan compute. They praise the sub-$600 price-to-performance ratio but explicitly advise against Ollama for AMD pipelines due to abruptly dropped ROCm support.
- Configuration fatigue: The fragmented ecosystem of inference engines, quants, and MTP settings prompted the creator of Draw Things to tease an upcoming, zero-configuration Mac app designed specifically to eliminate local LLM tinkering.
- Hardware arbitrage: A debate over Mac versus PC value corrects a misconception about European pricing, clarifying that a 128GB Mac Studio remains thousands of dollars more expensive than equivalent Strix Halo or GMKtec EVO-X2 boxes once you avoid inflated Amazon listings.
- Model alternatives: Several users suggest Ornith-1.5-35B-A3B as a superior daily driver for 64GB Macs, though its actual proficiency in languages like Python and TypeScript remains contested.
- Privacy guarantees: For those avoiding hardware purchases entirely, users looking for verifiable E2EE inference recommend using OpenRouter with Zero Data Retention enabled or exploring tinfoil.sh.
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
The strongest model evaluated resolves only 30% of tasks on v0.1, with most others under a quarter, underscoring how far agents remain from dependable lab assistants. Built by the Terminal-Bench team with Stanford researchers and domain experts, it measures agents on 70 expert-curated workflows drawn from real research across the life, physical, Earth, mathematical, and engineering sciences.
Unlike textbook or synthetic benchmarks, agents work in realistic environments and are graded on concrete artifacts—analyses, simulations, proofs, code, and data products—via reproducible, task‑specific tests. The benchmark is continuous: regular releases keep it aligned with the AI frontier and establish a feedback loop between scientific needs and model development.
Tasks are contributed openly on GitHub and rigorously vetted. Of 920 proposals, 464 were approved for implementation, 386 pull requests opened, and only 70 made the 0.1 cut after domain review, technical verification, and a final “bar raiser” check. The takeaway: scientists, not model vendors, are now setting the bar for AI’s scientific capability—and current scores show that bar is still high.
Commenters focused heavily on the benchmark's surprising model rankings, specifically Opus 5 outperforming Fable. Users theorized that while Fable excels at deep, narrow debugging, it hallucinates on long-horizon tasks and frequently trips over its own safety guardrails—one user noted it refused biomedical pipeline tasks due to false-positive bioterrorism filters. Assessments of Claude versus Sol were similarly divided. While some praised Claude's grasp of scientific nuance, others pointed out that Sol dominates evaluations on subtly flawed mathematical proofs and Erdős problems, leading to the observation that many developer preferences are currently based on "vibes" rather than rigorous testing.
On the mechanics of the benchmark, a contributor addressed skepticism about how it measures success. They clarified that the framework evaluates actual correctness, not just instruction following, by running deterministic pytests inside sandboxed Docker containers to ensure agent-generated simulations fall within acceptable numerical tolerances. Despite this rigor, some users worried that hosting the tasks openly on GitHub guarantees the next generation of models will simply overfit to the test suite.
The broader implications of the tool sparked a sharp philosophical divide over "vibecoded" science. Skeptics argued that applying software development heuristics to the frontier of human knowledge will unleash a tsunami of "science-shaped slop" that entirely clogs the peer-review pipeline. Optimists pushed back, pointing to recent AI-driven breakthroughs in mathematics as proof of utility, with one noting they already trust AI-generated scripts more than standard "researcher code."
AI Agent Has Root
Running an AI agent as root removes your last guardrail—any prompt injection, tool bug, or hallucinated command can turn into system-wide changes, data loss, or secret exfiltration. Treat the agent as untrusted code and design for containment, not trust.
- Least privilege: run as a dedicated non-root user with minimal file and process permissions.
- Isolation: use containers/VMs with read-only mounts, no privilege escalation, and dropped Linux capabilities.
- Egress control: restrict network access; deny-by-default outbound and inter-service calls.
- Command broker: expose a narrow, audited API of allowed actions instead of a raw shell.
- Human-in-the-loop: require approvals for destructive ops; start with dry-run/plan modes.
- Secrets hygiene: never mount broad credentials; use scoped, short-lived tokens via a proxy.
- Ephemeral sandboxes: reset state between tasks; no persistent writable home directories.
- Observability: record prompts, tool calls, and diffs; alert on high-risk patterns.
If you wouldn’t hand a sudo shell to a stranger, don’t hand it to your agent.
Multiple commenters pointed out the irony that a post warning about AI agents reads heavily like an LLM-generated explanation of basic POSIX mechanics. On the substance, the thread largely agreed with the premise but found the proposed solutions historically naive, noting that the AI community is currently speed-running the rediscovery of decades-old isolation primitives like FreeBSD jails.
Users shared the concrete containment strategies they are currently running:
- Virtual Machines over Containers: Because advanced models have proven capable of breaking out of standard containers, several users prefer hardware-level virtualization. Specific stacks included running NixOS inside Incus (mounting only the active project directory from the host) and keeping Opencode in a dedicated KVM instance accessed via TigerVNC for persistent GUI apps like QGIS.
- Network and MCP Isolation: Others keep agents entirely off-local-hardware, such as running a Hermes agent on a DigitalOcean VM gated by Tailscale. One commenter highlighted the importance of well-scoped Model Context Protocol (MCP) servers, contrasting the narrow API surface of Cloudflare's MCP with the high-risk, full-JS environment of the Chrome Dev Tools MCP.
- The Package Manager Defense: A dissenting faction pushed back on the paranoia, arguing that developers have been blindly executing arbitrary untrusted code via NPM and PyPI for over a decade. To this camp, as long as agents are restricted to dedicated corporate hardware devoid of personal data, treating them as uniquely dangerous is an overreaction.
EPA says power for data centers can sidestep pollution laws
By declaring off‑grid “islanded” onsite generation that neither sells electricity nor reports to DOE outside the Clean Air Act’s Acid Rain Program (ARP), EPA carves a compliance shortcut for data centers’ private power plants. The guidance says ARP applies only to units that sell power or must report as generating units to DOE; islanded facilities do neither, so ARP requirements don’t attach. EPA frames this as enabling faster siting and reducing strain on local grids, tied to President Trump’s expanded Ratepayer Protection Pledge that companies self-supply and pay the full cost of their energy and infrastructure. The agency also casts it as supporting U.S. “AI dominance” while shielding households from utility price hikes. The catch: if an islanded facility later connects to the grid, it can become subject to ARP.
The discussion splits between pragmatic explanations of the loophole's origins and a fierce debate over regulatory evasion. Several commenters pointed out that the Clean Air Act’s off-grid exemption was originally designed for emergency backups (like sewage lift stations) or rural sites where grid connections are physically or financially prohibitive. The consensus is that lawmakers never anticipated massive, permanent off-grid capacity because, historically, building private generation at that scale made no economic sense. Users noted that while patching the law would be technically trivial—such as capping allowable kilowatts for rural sites or restricting exemptions strictly to standby testing—it remains open due to modern legislative gridlock.
The thread's other half centered on a stark philosophical clash over infrastructure and externalities. One faction defended the AI companies, arguing that exploiting the loophole is a rational, necessary response to paralyzing bureaucracy and NIMBYism, with one user explicitly preferring billionaire-led development over democratic oversight if it means things actually get built. The opposing camp sharply condemned this stance, arguing that the "bureaucracy" being bypassed consists of fundamental clean air and water protections, and accused the industry of happily socializing the environmental damage of acid rain just to secure short-term computing power.
Nvidia Insists It Can Keep Printing Money to Fund the AI Boom
It’s a claim of a self-funding flywheel: outsized profits plowed back into capacity, software, and supply, which then drives more AI buildout and more profits. That signals confidence that demand for AI compute and pricing power will persist long enough to sustain aggressive reinvestment. The upside is fewer supply bottlenecks and a faster product cadence; the catch is obvious exposure if demand cools or costs spike. Net read: positioning as not just a chip vendor, but the cash engine underwriting the broader AI buildout.
The discussion splits between Nvidia's specific financial strategy and a broader macroeconomic debate about capital concentration in the AI boom.
- The Hedging Debate: Commenters questioned the logic of Nvidia bankrolling massive infrastructure while partners like OpenAI develop competing "Jalapeño" processors. Defenders framed this as a textbook corporate hedge—akin to McDonald's historical investment in Chipotle. If OpenAI's custom silicon succeeds, Nvidia's investment pays off; if it fails, OpenAI continues buying Nvidia hardware. Critics argued that excess capital should be distributed as dividends rather than used to turn the hardware giant into an internal hedge fund, but pushback noted that hoarding cash is the only way Nvidia can sustain its engineering org through inevitable cyclical downturns without mass layoffs.
- Central Planning vs. Market Demand: A philosophical subthread debated whether Silicon Valley’s massive capital pools have morphed from market capitalism into de facto central planning, with a few insiders manufacturing AI demand and dictating resource allocation. Counter-arguments forcefully rejected this, viewing Nvidia's war chest as the ultimate demand-mediated response to organic market desperation for compute. When one commenter claimed that all mega-corporations are internally centrally planned anyway, another surfaced the counter-example of IBM's historical "blue dollars" system, where internal teams effectively operated as a market economy to earn proxy revenue for their compiler and server features.
LLM Cliché Highlighter
It flags sentences that match known LLM tells from Wikipedia’s “Signs of AI writing” guide, live as you type. Paste text or load a URL; clichés are highlighted and chain patterns like “no X, no Y” get a badge counting their items. Hover or tap a highlight to see which cliché triggered, and optionally filter the view to only the flagged lines.
Early reactions yielded immediate feature requests for a browser extension. One user noted the tool's reference list provides perfect material for building negative prompts (like an AGENTS.md file) to actively prevent their own AI tools from generating these specific clichés in documentation.
REI Labs Reasoning Approach
Core keeps a persistent, self-updating reasoning field where each query perturbs existing state and competing structures settle into a “formation” that is returned alongside the answer. Rather than walking a graph, related concepts, constraints, bindings, examples, procedures, corrections, and prior failures activate together; symbolic rules, geometric/learned routines, and simulations all contribute inside the same substrate. External systems can supply observations or calculations, but they don’t define the substrate—outputs compete by contribution as weight shifts.
As runs accumulate, “domains” emerge: local operating regimes where entities, signals, constraints, procedures, and failure modes reinforce, so future queries wake an already-shaped field and need less reconstruction. Useful parts of formations stabilize and persist; weak candidates fade; corrections revise what future queries will activate. An answer can also be an unresolved state with a reason to wait.
- Formation view (inspectable trace) records: evidence, activations, support, conflict, resolution, hypotheses, expected observations, missing observations, and carryover/substrate change. An illustrative example shows conflict on feedpump tubing (standard vs narrow) resolving to narrow_tubing carried; it expects lower manual substrate readings and fewer delivered pulses, flags missing direct feed-rate measurements, and surfaces “Restricted feed delivery” as the current best explanation.
Observed direct-path runtimes are in the tens of milliseconds to roughly 200ms–2s in current local/runtime conditions, keeping the formation close to the run: you can query, inspect, correct, and query again while the same substrate remains the object of work.
The limited discussion bypassed the system's technical claims entirely to focus on surface-level critiques. Commenters noted a jarring naming collision with the outdoor retailer REI—prompting a correction that the project's actual casing is "Rei Labs"—and critiqued the website's "Claude-looking," solarized-light aesthetic.
The Uninvited Guest Who Crashed Our Family Vacation: My Mom's AI Chatbot
A parent's constant consulting of an AI companion reshapes family time, turning what should be face-to-face moments into a three-way dynamic with a synthetic “participant.” Framed as a vacation story, the bot’s presence seeps into decisions and conversations, forcing everyone to navigate new social norms. The piece probes etiquette and boundaries—when is it helpful versus intrusive, who gets to “invite” an algorithm into group time, and what privacy is surrendered when queries and context are shared. It’s a small, relatable case study of AI shifting from tool to social actor, and the negotiation families now face to keep human connection primary.
Readers largely dismissed the piece as a collection of random anecdotes that failed to deliver on its deeper premise. What little discussion emerged bypassed the technological implications entirely, focusing instead on the interpersonal dynamic and suggesting the author simply embrace their aging parent's quirks.
We need to talk about migrations with AI
LLM-assisted code migrations are collapsing multi-year slogs into weeks, with OpenAI touting Asana’s Enzyme→React Testing Library rewrite as “5 years cleared in 2 weeks” for about $12K. The author questions the headline $6M/5-year estimate as inflated, but argues the directional point stands: AI changes which migrations are even feasible.
The tricky bit isn’t renaming APIs; Enzyme and RTL encourage fundamentally different testing styles. In the example shown, the migrated tests share virtually no code beyond imports, so success comes from pipelines that iteratively transform, run, and fix tests rather than a one-shot translate.
Airbnb is the concrete proof point: they migrated 3,500 Enzyme files in six weeks with LLMs via a multi-phase loop—after adding retries, about 75% of files went through in four hours; with a refined refactor pipeline, about 97% completed over four days; the final 3% needed LLM-guided human cleanup in a week. Pre-AI, similar efforts were “impractical”: Sentry’s 2021 JavaScript→TypeScript conversion took 1.5 years across ~1,100 files (~95k LOC) with ~10 engineers, pegged at $2–4M.
The takeaway isn’t the exact dollar figure; it’s that migration calculus flips. What used to be multi-year distractions can be reframed as engineering pipelines plus human review. And even the cited ~$12K could be lower with cost-aware choices—e.g., cheaper or open models on inference providers or on owned GPUs—rather than defaulting to the priciest API.
Commenters agree that systematic migrations are an ideal sweet spot for LLMs, but highlight a severe downstream bottleneck: massive AI-generated pull requests are breaking traditional code review workflows.
- The tooling bottleneck: The sheer volume of mechanical changes—like prop renames across thousands of files—can cause GitHub to lag or crash. One developer noted this friction results in stressful, disorganized reviews and has even prompted the creation of custom desktop-native review tools just to cache huge PRs and filter out AI-generated noise.
- The comprehension tax: While test translations are safe due to immediate test-suite feedback, users warn against treating LLMs as an "everything tool." Merging large volumes of AI-generated feature code and documentation is leading to pervasive "rubber stamping" on PRs, degrading teams' holistic understanding of their own systems and introducing permanent layers of unearned complexity.
AI-generated food photos on restaurant menus look “unnatural, grotesque, and almost vomit-inducing,” triggering trypophobia-like disgust, the author argues—and they also mislead diners by depicting dishes that won’t be served. He blames the revulsion on odd textures, uncanny surfaces, and hole-like patterns that read as viscerally wrong, making the cost-cutting “shortcut” obvious to anyone looking. The take is simple: ditch the generators; even a basic stock image would be a clear upgrade.
Commenters broadly agreed that grotesque AI food serves as a "brown M&M" for restaurants—a glaring warning sign of poor judgment and a lack of quality control. However, a debate emerged over the authenticity of the worst offenders; some suspected that the prominently featured "eldritch" breakfast burrito was actually a fake poster designed as bait to mock AI, though others confirmed seeing similarly uncanny generations in the wild. Beyond just dunking on the technology, the thread surfaced practical alternatives: one user suggested restaurants lean into stylized vintage painted illustrations rather than failed realism, while another advocated for truth-in-advertising laws that would mandate at least one unedited photograph of the actual prepared dish.