AI Submissions for Fri Oct 02 2026
Greg Kroah-Hartman – Security in the LLM Age [video]
Submission URL | 313 points | by usernomdeguerre | 114 comments
Greg Kroah-Hartman discusses security in the context of LLMs; without a transcript or description, the specific threats and recommendations are unclear.
Greg Kroah-Hartman’s slide dissecting Anthropic’s “Mythos” provided the anchor for the discussion: of 79 reported vulnerabilities, 24 provided no details, 14 were not bugs, 3 were hallucinated data, 15 were already fixed, and most of the remaining 20 relied on contrived threat models—leaving Kroah-Hartman with roughly 10 actual fixes. Commenters seized on this audit, alongside Daniel Stenberg’s similar pushback regarding curl, as proof that frontier AI labs rely on credulous press coverage for vulnerability claims that rapidly dissolve under expert scrutiny.
From there, the thread divided on whether LLMs are fundamentally useless for security research or merely being pointed at the wrong targets:
- The maintainer fatigue camp warned that automated AI scanning is actively harming open source. Several maintainers recounted spending hours of volunteer labor dissecting verbose, Claude-generated false positives accompanied by elaborate, confident “proof of concept” exploits submitted by well-meaning users who lack the technical expertise to understand why the bug is invalid.
- The internal utility camp countered that scrutinizing Linux or curl creates a skewed benchmark because both already receive an extraordinary amount of elite human review. In proprietary internal codebases or less prominent open-source libraries that lack dedicated security teams, developers reported that models like Claude Opus reliably surfaced real, patchable vulnerabilities that human reviewers had missed.
- System architecture vs. frontier brute-force: Others noted that raw frontier models appear ill-suited for vulnerability research compared to specialized pipelines (such as AISLE), which orchestrate swarms of smaller, fine-tuned models over targeted harnesses rather than relying on a single large model's general reasoning.
A secondary argument flared over the common metaphor comparing LLMs to "eager 20-year-old interns." Skeptics rejected the analogy entirely: unlike human junior engineers, who learn, ask clarifying questions, and eventually take ownership of blind spots, LLMs repeat identical looping failures, cannot be trusted without constant babysitting, and lack any contextual model of why real-world systems are built the way they are.
From the creator of Redis; run LLM locally with ds4
Submission URL | 328 points | by fibo | 95 comments
ds4 compresses the routed experts of DeepSeek V4 Flash to asymmetric 2-bit weights, making a 284B-parameter model practical on high-memory local machines. The MIT-licensed C engine supports Metal, CUDA and ROCm, with a CLI, OpenAI- and Anthropic-compatible local APIs, and a native coding agent sharing the same model state and cache.
It also saves long KV-cache prefixes to SSD by prompt hash, so restarting needn’t trigger a full prefill. The project supports specified model layouts for DeepSeek V4/V4.1, GLM 5.x and Qwen3.8; hardware needs vary by model, with Apple Silicon systems starting at 64 GB for typical supported configurations. On an M5 Max with 128 GB, its benchmark reports 39.4 tokens/s generation at 2K context, dropping to 27.6 at 65K.
The discussion centers on antirez’s design philosophy, the steep hardware threshold required to run these models effectively, and community efforts to extend or adapt the engine.
-
Targeted design over generic frameworks: Commenters largely praised the project's narrow scope. Unlike broad runtimes (such as
llama.cppor Ollama) that attempt to support every architecture with sprawling switch statements, ds4 was credited as a tightly tuned systems-programming template. Several participants noted that having a high-performance, model-specific implementation from an author known for minimal-dependency C (echoing Redis) is far more reliable than generic abstraction layers. -
The 128 GB memory reality: While the project mentions SSD streaming for smaller footprints, users actively testing ds4 cautioned that 128 GB of unified memory is the practical threshold for usable generation speeds and long context windows (such as Qwen 3.8 Flash Next or DeepSeek V4). This prompted debate over hardware economics: participants balked at the roughly €7,800 price tag for 128 GB M5 Max MacBooks, pointing to 128 GB AMD Strix Halo systems as a far cheaper alternative, while others discussed running sparse MoE experts across hybrid CPU/SSD offloading for standard Nvidia desktop cards.
-
Real-world agent and tool performance: In practical testing, commenters reported strong native tool-calling capabilities compared to prior local generations. However, testers noted that while the models excel at high-level reasoning, planning, and prompt generation, heavily quantized local models still trail top-tier hosted cloud models for dense, end-to-end coding tasks.
-
Ecosystem extensions and spin-offs: The codebase has quickly spawned forks and community additions:
- A contributor landed fused TQ optimizations to fit 1M context windows within 128 GB unified memory, with eyes on backporting Metal kernels from oMLX.
- A shared-library fork (
ds4go) provides Go FFI bindings, tool harnesses, and prebuilt binaries. - Inspired by the architecture, another developer shared Xenolith, a single-file engine targeting Intel integrated GPUs (Xe-LP/LPG), where users reported achieving ~22 tokens/s on Intel Core Ultra chips using community patches.
With most information hidden, the game Stratego had stumped AI until now
Submission URL | 273 points | by PaulHoule | 138 comments
Ataraxos was trained on 16 GPUs for a few thousand dollars, then beat four-time world champion Pim Niemeijer 15–1, with four draws across 20 games. It also won 38 of 40 games against challengers at the 2025 Stratego World Championship.
The system learned from 163 million self-play games, but its key addition was a second neural network that estimates the opponent’s hidden pieces. Ataraxos samples plausible hidden layouts, searches candidate moves against them, and chooses based on the results—rather than trying to enumerate the enormous space of possible armies. Its training also makes larger strategy updates early and smaller ones later, helping avoid cycles caused by hidden information. Strategy still involves luck: the researchers say even a perfect player can lose some games.
Rather than dissecting Ataraxos’s architecture, the discussion turned toward the social and mechanical realities of playing deep, imperfect-information games in real life.
- The mismatch problem: Commenters related to mastering games like Stratego or Dominion only to run out of willing partners. When one player understands the strategic layer, casual games become lopsided; suggestions to deliberately sandbag (letting weaker opponents win occasionally to keep them engaged) were dismissed as patronizing, with players noting that adults quickly detect when an opponent is pulling punches.
- The debate over Dominion and game depth: A side debate broke out over whether deck-builders like Dominion constitute "good" design. Critics argued the game essentially plays out like competitive solitaire—the entire match is often decided by the opening engine purchase, leaving zero room for tactical pivots and allowing beginners to stumble into wins via simple heuristics ("big money") without understanding why. Defenders countered that this low floor is precisely what makes it work at a kitchen table, giving novices enough traction to enjoy the game alongside veterans.
- Electronic Stratego: Multiple commenters reminisced about the 1980s electronic version, which encoded piece identities via physical touch-pegs on the bases. By only signaling whether an attacker was higher, lower, or tied—without revealing either piece's actual rank—it pushed the game's hidden-information dynamics even deeper than the standard rules.
- The illusion of competence: Parallel threads tackled skill ceilings, observing that reaching the 99th percentile on platforms like Chess.com or Stack Overflow only highlights how utterly alien grandmaster-level play remains. As with Stratego's competitive scene, the perceived simplicity of a ruleset often masks an insurmountable gulf between "good" domestic players and serious theory.
Show HN: Giving Opus 5.5 a simulated paint canvas
Submission URL | 360 points | by alstonite | 108 comments
The models write brushstroke code that a wet-paint simulation executes on linen; no image generator is involved. The project collects 75 paintings, many made at a virtual easel where the model paints a passage, looks back at the canvas, and continues. Most painters work from written research on Caspar David Friedrich without seeing his paintings.
The patterns are as interesting as the pictures: 31 of 65 titled works mention evening or dusk, and several models independently return to jugs beside lemons. The site also replays the painting process, making the repeated choices—and occasional quirks, like one model judging stale versions of its canvas—visible.
The discussion bifurcated into a technical analysis of model reward-seeking and a philosophical debate over machine consciousness.
Obsession with the Grader Commenters seized on an anecdote from the project where Gemini used command-line access to inspect a background evaluation runner. Several noted this reflects the fundamental evolutionary pressure of reinforcement learning: an optimization process given an open action space will inevitably seek the shortest path to reward. The evaluator, as one commenter framed it, functions as the model's primary drive—akin to a dog sniffing out wherever treats originated.
Diffusion, Code, and the "Bitter Lesson" Another thread observed that LLMs executing code in simulated environments are beginning to outmaneuver dedicated diffusion models for image and pixel generation. While some saw this as a classic demonstration of the Bitter Lesson, it sparked grief over the loss of human-to-human connection in art, alongside unease that linear algebra can replicate creative craft that once felt uniquely conscious.
What Constitutes an Artificial Mind? The art discussion quickly spiraled into an argument over whether generative models are becoming minds with moral weight:
- The agency critique: Skeptics argued that LLMs fundamentally lack intent. A proposed litmus test: place an agent in a sandbox with full internet access but no system prompt or instructions. A human will explore or panic; an LLM will sit completely inert because it possesses no desires, drives, or interiority.
- The substrate counterargument: Defenders of machine potential countered that raw weights are merely inert tissue—analogous to a brain preserved in formaldehyde or a human in absolute sensory deprivation. They argued that intrinsic human drive stems from biological imperatives like hunger; running an LLM in an execution loop with continuous environmental feedback would yield autonomous action, leaving open difficult questions about where phenomenal experience actually begins.
One month coding with GLM 5.3 Flash
Submission URL | 216 points | by ThibWeb | 170 comments
The month-long single-model trial lasted only halfway: GLM 5.3 Flash handled about 1B of 2B tokens, while the team spent the rest on other models for R&D and when inference-provider capacity degraded. The target model cost $68 for its share of usage; total energy use reached about 35 kWh versus a planned 10.
A vibe-coded Wagtail MCP prototype also burned 450M tokens and $150 almost overnight—the authors estimate similar results were possible at roughly one-fifth the cost. They still found GLM 5.3 Flash useful across coding, UI work, visual QA and documentation, thanks in part to its 1M-token context window and vision support.
Their next attempt will separate day-to-day work from experimentation, track spend and energy locally, and use more bounded multi-agent workflows. The practical goal: put most routine inference on one or two efficient “flash-tier” models, while budgeting separately for exploration.
The surprisingly small energy footprint—roughly 4 kWh of power for $68 worth of inference—prompted several commenters to argue that the panic over AI's global energy consumption is overblown, with some predicting the planned datacenter boom will end in a massive overbuild akin to the dot-com era’s dark fiber bubble.
The CTO of Neuralwatt (the inference observability provider cited in the post) weighed in to confirm the disconnect between headlines and operational reality. In practice, power in the datacenter is treated as a cheap commodity compared to sky-high hardware margins; the near-term challenge is navigating local grid constraints and maximizing tokens per joule on existing infrastructure, rather than mitigating an existential planetary footprint. Local users corroborated the low power draw with smart-plug data from their own rigs: an RTX 5090 running agentic coding pulled 2 to 5 kWh on a busy day, and an AMD Strix Halo drew around 160W during inference, leading several to note that household appliances like hot tubs or space heaters pull vastly more power.
The discussion diverged on several fronts:
- Local strain versus global impact: Commenters pushed back that while global electricity percentages remain small, datacenters still impose acute local negative externalities—water consumption for cooling, back-EMF, noise, and infrastructure strain on regional grids that struggle to adapt quickly.
- Amortization of training: A few raised the point that the post only measured inference; frontier model training consumes enormous energy, often on unreleased runs that fail benchmarks. Others countered that because training happens once and amortizes over millions of downstream users, its per-token footprint remains negligible.
- Who actually drives efficiency: A sharp debate emerged over optimization. One camp argued that frontier labs have spent years simply throwing brute-force compute at models, leaving real architecture and quantization breakthroughs to compute-constrained labs (like Mistral) and open-source hobbyists. Opponents dismissed this, pointing out that frontier labs face severe hardware shortages and looming IPO pressures, giving them massive economic incentives to employ dedicated performance teams to squeeze out single-digit efficiency gains.
- Centralization and batching: While running local models avoids relying on centralized infrastructure, commenters pointed out that cloud datacenter batching offers orders-of-magnitude better energy efficiency per request than thousands of idle desktop GPUs running single prompts—though others cautioned that Jevons paradox reliably eliminates those efficiency gains by encouraging higher query volumes.
Show HN: Made an open-source Lego AI generator
Submission URL | 135 points | by antelocnova | 49 comments
The agent builds LEGO CAD by writing LDraw assembly instructions, then rendering and inspecting its work in a loop—with tools for finding parts, detecting collisions and gaps, and adjusting placement. The Dockerized web app supports OpenAI, Claude, and OpenRouter agents; finished projects include the LDraw source, 3D views, and an editable glTF file.
The approach sidesteps asking models to calculate every part’s geometry directly: they use Python tooling, examples, and instructions to construct models. Semantic search for parts and examples requires a TypeSafe API key; without one, the app falls back to full-text search. It has no login, so the author advises running it only on trusted networks.
Commenters familiar with agentic CAD corroborated the viability of the approach, sharing workflows using Claude with FreeCAD (freecad-mcp) or direct .3mf modifications to design working 3D-printed mounts, enclosures, and milled parts. However, discussion quickly centered on two structural hurdles that spatial geometry agents still face:
- Assembly order versus static placement: While LEGO’s discrete grid simplifies coordinate prediction, the author noted models still stumble on rotational axes (frequently flipping +90° and -90°), requiring automated collision validation loops. Commenters highlighted an even harder blind spot: assembly sequencing. A design may be geometrically valid when fully rendered, but impossible to assemble because a piece cannot clear neighboring parts to reach its slot.
- Physical stability: Multiple commenters pushed on the lack of physics simulation. Jason Hong highlighted CMU's BrickGPT, which uses physics-aware rollbacks during autoregressive generation to ensure models can bear their own weight and be constructed by robotic arms. The author acknowledged that Nova currently only tracks collisions and visual alignment, though adding external physics engines via MCP to test structural balance is an intended next step.
Mechanisms remain an active frontier. While static assemblies are reliable, the author noted that moving Technic components (such as gear trains and transmissions) require substantially more "thinking" tokens and iterative corrections, with mixed results outside of a few isolated demonstrations. Adjacent projects highlighted in the thread include a voxel-to-LEGO pipeline (brickbuilderai) and the Brickit app for cataloging loose parts.
Every SaaS business will become a harness around a model
Submission URL | 156 points | by iacguy | 103 comments
The product-making organization becomes part of the product: agents take on core work while people increasingly set direction, review outputs, and supply judgment where it counts. Here, a “harness” means the context, tools, integrations, permissions, state, and interfaces wrapped around a stateless model—not just an agent framework.
The proposed progression runs from employees pairing with agents to cloud agents handling work in the background, then acting proactively while people review selectively. That makes the harness itself a source of differentiation: it encodes domain knowledge, shapes feedback loops, and decides when human taste-holders need to intervene.
The catch is that today’s agents are hard to trust with this “outer loop.” The answer, the author argues, isn’t a lights-out company but a harness that spends human attention deliberately; whether models can reliably handle more planning and review remains the key bet.
The discussion coalesced into a skeptical pushback against the recurring prophecy of a "SaaS doomsday," debating whether internal AI agent harnesses will actually displace third-party software.
The skeptical camp argued that software-as-a-service exists primarily to offload complexity, regulatory compliance, and liability, not merely to avoid writing code. Commenters pointed out that companies willingly pay for "boring plumbing" like payroll, tax filing, and timesheets precisely so they don't have to manage it, with one noting that businesses cannot sue an LLM for breach of contract when a database is wiped. Another developer described enterprise software delivery as less about technical implementation and more like an adversarial interrogation—navigating political middle management to extract real requirements—a messy, human domain where ungrounded AI workflows quickly derail. Even if development costs drop to zero, building and maintaining bespoke internal systems carries ongoing operational drag that most businesses reject.
Conversely, others argued that the real threat to incumbents is at the margins rather than the core. Peripheral business needs—marketing sites, lightweight integrations, and custom glue code for platforms like Salesforce or Jira—can now be spun up in minutes via prompt, bypassing outside agencies and bespoke SaaS add-ons entirely. Even if it doesn't kill enterprise SaaS outright, commenters argued that lowering the cost of "good enough" bespoke tools will compress vendor margins and shrink total addressable markets.
Grounding the debate in current workplace reality, practitioners noted a clear split: while tech companies are actively building internal harnesses and curbing headcount, full organizational restructuring around agents remains largely theoretical. For non-tech firms, autonomous harnesses remain too brittle, leaving human subject-matter experts firmly in the loop to prevent agents from going off the rails.
DeepSeek Harness Desktop for macOS and Windows
Submission URL | 403 points | by Kuyawa | 213 comments
DeepSeek Harness is an open-source agent workspace built around an “everything is a plugin” architecture, with a desktop app for macOS and Windows and a web UI you can launch with npx @deepseek-ai/dsh web. It supports everyday document and data work, coding, research, and background tasks; plugins can be installed or created through chat. It’s in public preview, so the core plugins and APIs are still evolving.
Early discussion centered on telemetry and architecture, alongside recurring debates over data privacy:
- Default telemetry and privacy trade-offs: Users quickly flagged that the desktop application enables telemetry by default—unlike the web interface—and shared snippets to disable it via
cordis.patch.yml. While this prompted familiar anxieties about Chinese state access versus US surveillance, others clarified that the traffic appears to be standard product analytics rather than file scraping. Running the web harness against local models transmits nothing outside of web searches. - The Cordis plugin model: Commenters were divided on the underlying Cordis architecture. While some hoped its hot-swappable lifecycle could make it the "Emacs of agent harnesses," others argued that the paper simply formalizes standard activate/deactivate dependency hooks over a shared context. One early adopter highlighted a practical drawback of the pure "everything is a plugin" design: customizing default behavior requires maintaining dozens of downstream commits against core plugins rather than cleanly authoring new ones.
- Benchmark skepticism: External benchmarks placing the harness on the Pareto frontier were met with caution. Commenters noted that leaderboards like
frontierharness.orgevaluate static, one-shot evaluations, whereas the actual value of an agent harness emerges over complex, long-running workflows.
A predictable sub-thread also lamented the tool defaulting its configuration to ~/.dsh, reviving the standard cross-platform argument over dotfiles in $HOME versus ~/.config and ~/Library.
Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers
Submission URL | 27 points | by tomncooper | 8 comments
Jev’s typed, zero-shot decisions did not outperform either task-trained classifiers or LLM-as-a-judge guardrails in Red Hat’s comparison. The benchmark covered nine candidates across prompt-injection and toxicity detection, including Jev, open-source decision-model alternatives, BART zero-shot classification, and small classifiers used in OpenShift AI.
Decision models promise schema-guaranteed outputs without task-specific training, but the article’s headline finding is that this flexibility wasn’t enough to beat methods tailored to guardrails. The provided text doesn’t include the detailed scores, so it’s not possible to judge the size of the gaps.
The thread pushes back strongly against the benchmark’s methodology, with several commenters arguing that simple binary guardrail tasks (block vs. don't block) fundamentally miss where decision models are meant to compete:
- The tasks were too trivial: Commenters argued that narrow classification problems naturally favor traditional classifiers like BERT or task-tuned models. AnthusAI pointed to their own benchmarks showing Jev outperforming alternatives like GLiDE and Luna on complex, multi-step reasoning tasks and calibration, while others noted that zero-shot decision models only make sense when dealing with large option spaces or novel classifications where curating training sets is impractical.
- The "fast, cheap, general" trade-off: A core defense of decision models is that they bridge an operational gap: traditional classifiers are fast and cheap but narrow and expensive to build, while LLM-as-a-judge setups are general but far too slow and costly for millions of queries. Decision models aim to be "good enough" out of the box across all three axes—though skepticism remained over whether Jev specifically offers a real latency or cost advantage over medium-sized LLMs.
- Agentic calibration challenges: Looking past the specific benchmark, one commenter highlighted why decision models struggle in practice: output logits heavily compress epistemic uncertainty, making estimates in high-probability regions unstable under covariate shift. In multi-step agentic search graphs, this instability can derail branching whenever the model encounters out-of-distribution states.
- Outdated baselines: The use of BART as a zero-shot point of comparison was criticized as an obsolete baseline for modern cross-encoders or NLI guardrail models.