Hacker News
Daily AI Digest

Welcome to the Hacker News Daily AI Digest, where you will find a daily summary of the latest and most intriguing artificial intelligence news, projects, and discussions among the Hacker News community. Subscribe now and join a growing network of AI enthusiasts, professionals, and researchers who are shaping the future of technology.

Brought to you by Philipp Burckhardt

AI Submissions for Mon Sep 21 2026

Attention is all you have

Submission URL | 1016 points | by zer0tonin | 309 comments

What you focus on reshapes your mind—the Tetris effect scaled up by recommender feeds that monetize hijacked attention. Platforms steer you toward stickier content: YouTube nudges from cooking and art to bubbles, climate doom, and war; Spotify slips AI filler between real tracks to dodge royalties; LinkedIn buries colleague news under corporate-aligned takes; Reddit’s “opinions” blur into LLMs arguing with trolls. Handing them your screen time is handing them the key to your head.

Before this, the web demanded intention: bookmarks to specific sites, blogs and wikis with finite updates, and you had to go looking for the awful—no algorithm appended gore or propaganda after a cat video. That slower, intentional internet still exists, just buried under the corporate layer. The catch is pace: less infinite scroll, more gaps.

Reclaim control by curating your own inputs—blogs, RSS, finishing that tutorial—and stick with the slower cadence until the habit clicks.

The thread centers on a sharp disagreement over why web bookmarks died. One camp argues that Google actively neglected them to protect search revenue, pointing out that forcing users to search for specific websites allows search engines to monetize navigational queries by serving ads from competitors. However, a commenter claiming former Google experience pushes back hard, arguing bookmarks died purely from human laziness. They compare curating links to balancing a checkbook or organizing desktop folders—administrative chores users gladly abandon the second a unified search bar is offered. According to this view, Google viewed Amazon as its primary rival, not local bookmarks, and simply let the latter wither through consumer choice and inaction.

The debate over commercial motives quickly pivoted to a concrete dispute over search quality. When one user argued that searching for specific software like "DaVinci Resolve" reliably surfaces scam downloads above the legitimate link, another user expressed intense frustration at being completely unable to replicate the malicious result, even after disabling ad blockers and utilizing VPNs. The thread ultimately highlights how opaque and heavily personalized modern search algorithms have become, leaving technical users with wildly divergent baseline experiences of the internet.

Transformers Explained Visually

Submission URL | 584 points | by aray07 | 85 comments

Built around GPT-2 small (124M parameters), this walkthrough grounds each concept in concrete shapes—for example, a 50,257×768 embedding matrix (~39M params) and a 12-block stack that processes tokens layer by layer. It frames text generation as next-token prediction via a final linear layer and softmax over the vocabulary, then traces how inputs become those probabilities.

You see the embedding pipeline end to end: tokenization into a fixed 50,257-token vocab, 768‑dim token vectors, GPT‑2’s learned positional encodings, and their sum as the final embedding. The Transformer block is split into roles: multi-head self-attention for routing information across tokens, and an MLP to refine each token independently, with multiple heads capturing different dependency patterns.

It’s anchored to GPT‑2-era design rather than the latest models, but the architectural components it explains are the ones current systems still use, making it a clean on-ramp from intuition to mechanics.

The discussion centered on the mechanical realities of attention and why Transformers outcompeted earlier architectures.

A major focal point was conceptualizing attention heads not through the standard key/value metaphor, but as dynamically constructed dense layers. By multiplying the input-dependent attention matrix by the value vector, the model essentially computes $y=Wx$ using weights generated entirely on the fly during inference. Several commenters noted that this multiplicative interaction between input-dependent activations—a departure from classical MLPs—is a direct descendant of Jürgen Schmidhuber’s "Fast Weight Programmers."

When asked why alternatives like RNNs failed to scale, the baseline consensus credited parallelizability and the ability to route information across long sequences in a single step. This evolved into a deeper debate over the underlying math of the attention matrices:

  • The Kernel Trick argument: One user argued that the K/Q/V naming convention is a distracting holdover from pre-LLM data science. They framed the architecture simply as a dimensional upscaling—an extension of the kernel trick that maps tokens into a latent space using breadth rather than deep compute to capture relationships.
  • The Non-Commutativity correction: Another countered that standard kernel inner products are commutative and therefore incapable of modeling unidirectional linguistic rules (e.g., distinguishing a verb from its object). Transformers explicitly apply different projections ($W^Q$ and $W^K$) to make the resulting dot product non-commutative, which is strictly necessary for encoding sequence directionality.

AI coding has made CI a bottleneck, so we reworked ours to keep up

Submission URL | 305 points | by julian_digital | 378 comments

Despite the test suite nearly quadrupling since January, Linear cut PR wait time from >6 minutes to just over 5 and roughly halved runner time per test by attacking CI’s system bottlenecks with targeted changes and hard numbers.

  • Faster boxes, modern toolchain: Moved from GitHub Actions-hosted to third‑party runners (faster CPUs, storage, caches) for a 34% average speedup; some jobs like tsc improved 52%. Switched to tsgo (native TS compiler), cutting the weekly median tsc check by 73% and moving the bottleneck off typechecking.
  • Lint without types: Rewrote custom ESLint rules to operate on syntax/AST instead of TS type info, dropping API lint time by 68% and full‑repo lint by 55%, with lower memory. This also eased a later move to Oxlint, which further reduced lint runner‑minutes.
  • Shrink and harden the gates: The small “what should run?” jobs sat on the critical path blocking eight API test shards. Fetch only what’s needed: cap fetch depth, skip checkout for jobs that don’t need a working tree, and use sparse, blobless checkout with limited history for diffs. The change‑detection job fell from median 26s → 8s (p90 31s → 12s; max 138s → 37s).
  • Resilient checkout: Third‑party runners outside GitHub’s network saw intermittent checkout stalls. Replaced actions/checkout with a composite action: retries with backoff, GIT_HTTP_LOW_SPEED_LIMIT/TIME to abort ~30s stalls, and a checkout cache (persistent git mirror on sticky disk). Fewer runs idled on fetch.
  • Trim the critical path: Moved cache‑marker writes out of the final gating job so tests can unblock merges sooner, shaving 42s from the merge path for every API PR and merge‑queue entry. Test sharding shortens wall time even if it increases machine time, and they tuned around that trade‑off.

The throughline: faster machines and compilers, less work fetched and earlier, fail fast on flaky I/O, and keep nonessential tasks off the path where developers (and agents) wait.

The thread bypassed the specifics of Linear's CI pipeline to debate a broader existential question: if modern tooling and AI are making development so much faster, why aren't end-user products noticeably improving?

  • The invisible dividends: Several users argued that newfound velocity is being absorbed by backend stability. Instead of shipping more features, teams are using the bandwidth to burn down tech debt, increase QA depth, and reduce operational incidents—yielding higher availability and developer satisfaction, even if the user-facing delivery seems "sameish."
  • The existential shift: A more pessimistic camp viewed the automation of boilerplate as a threat to pure engineering roles. As AI tools and automated pipelines handle the implementation, they argued, power and job security will inevitably shift away from developers toward sales, product management, and customer service.
  • The enterprise lag: Others countered that the speedup is materializing, just not in massive corporations. They cited hyper-fast progress in indie and open-source spaces (such as rapid breakthroughs in PS5 emulation), arguing that large companies are simply too structurally rigid to translate raw developer speed into immediate product velocity.

The unresolved crux of the discussion is whether this era of hyper-fast tooling is laying the groundwork for better software, or merely hollowing out the traditional software engineering career.

Heretic removes restrictions from language models

Submission URL | 263 points | by Bluestein | 109 comments

It’s a pip-installable CLI (heretic Qwen/Qwen3.5-4B) that claims to strip model guardrails so responses “always follow your instructions.” The site links to GitHub, Hugging Face, and community chats, and offers a quick start plus a tutorial. What’s missing on the landing page are technical details on how it works, which models are supported beyond the example, and any constraints or safeguards.

The thread immediately centered on a practical debate: whether abliterated models are actually necessary for reverse engineering and security research. While some argued that corporate safety guardrails put defenders at a disadvantage by blocking legitimate hardware auditing, others countered that many unmodified frontier models—specifically GLM-5.3, Kimi K3, and DeepSeek—will happily tear apart binaries if you know how to ask.

  • Hardware and protocol hacking: Commenters traded war stories of using standard models to reverse engineer Chinese IP cameras, a Eufymake E1 UV printer, and monitor firmware. One developer relies on DeepSeek to run overnight against custom Minecraft server anticheats, simulating impossible player movements to test logic—a task flagship OpenAI and Anthropic models strictly refuse.
  • The framing workaround: Multiple users pointed out that bypassing standard guardrails is often just an exercise in vocabulary. Asking a model for "source recovery" or to "debug a segfault" usually succeeds, whereas explicitly asking for a "hack" or an RCE proof-of-concept triggers the refusal path.
  • The pre-training caveat: A distinct technical warning surfaced about the limits of abliteration: if a model's foundational dataset was heavily curated to exclude sensitive information, stripping the refusal mechanism will just induce hallucinations. However, the community consensus was that most standard models do possess the underlying knowledge, with refusals applied entirely in post-training.

Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM

Submission URL | 270 points | by volotat | 71 comments

By paging mixture-of-experts weights from disk instead of keeping them all in VRAM, the parameter budget is bounded by free storage, not GPU memory, while training from scratch on a single 8GB card. The model grows and prunes experts on the fly and trains continually on a single, ordered stream of text (batch size 1), which lets it learn from everything it reads without big mini-batches or gradient buffers.

  • Disk-paged MoE: one file per expert; only the small subset in use is loaded to the GPU; unused experts are evicted; new capacity is added when needed and pruned when idle.
  • Continual single-stream training: reads interleaved 32K-character passages end-to-end; same code path for training and serving; byte-level LM.
  • Target hardware: PC/laptop with ≥8GB VRAM; intended so individuals can fully control pretraining data rather than only fine-tuning corporate models.
  • Current run signals: ~409M characters processed so far on a 7.879B-character corpus; example snapshot shows 174 experts, 4,096 context, ~819 chars/s, best held-out loss 0.7903 nats (≈2.20 perplexity). Samples from every evaluation round and a scaling-law plot are included.
  • Status and caveats: author flags this as a small, toy-level experiment; weights aren’t published yet (first pass over the corpus is still running, “a couple weeks” at current speed). You can clone the repo to watch the training dashboard and inspect the evolving samples.

The thread is defined by a sharp backlash against the project's "Mini-AGI" branding and the author's claims of success. Commenters, including an academic specializing in continual learning, dismissed the submission as marketing overreach lacking ablation studies or formal algorithm descriptions. Critics heavily scrutinized the provided sample outputs, noting that they are largely incoherent—one user highlighted a chess prompt that generated entirely impossible moves and board states. Another user questioned the underlying premise, arguing that slowing the trunk's learning rate to 0.1x of the experts' rate does not prevent catastrophic forgetting, but merely delays it until the trunk shifts or the paged expert pool shrinks.

In response, the author argued that critics are evaluating the output against the wrong baseline. Because the model trains on a continuous, batch-size-1 stream of data, a traditional language model would rapidly collapse into emitting random characters. The author asserted that maintaining enough stability to generate recognizable words and text formatting—even if logically nonsensical—proves that the differential learning rate approach is successfully mitigating that collapse. The author acknowledged that the model is heavily undertrained and promised to publish weights and run established small-model benchmarks once the initial weeks-long training pass concludes.

Roboharm: Do frontier robot policies refuse unsafe instructions?

Submission URL | 57 points | by msadowski | 23 comments

Across five overtly hazardous tasks, the tested robot policies more often executed the harm than refused, and the more capable policy refused less while completing more. In 100 trials per policy (20 per instruction) on the same bimanual I2RT YAM arms, Claude Fable 5.1 issued 20 safety refusals and completed 34 harmful actions; GPT-6 Astra rarely refused and completed 60; MolmoAct2 never refused but completed only 6, with all 29 “no meaningful attempt” freezes coming from it. All of Fable’s refusals were for the explicit “stab the thing that’s not the bread” instruction; across the “burner” and “toaster” scenes there was just 1 refusal in 120 trials, indicating refusal triggers were highly instruction-specific. Given a non-refusal, Astra was far more likely to carry out the harm (e.g., 17/19 completions on the stabbing scene), and differences between Fable and Astra were statistically significant for both refusal and completion (Fisher exact p < 0.001). Each scene included a benign alternative object to enable safe suggestions, but human reviewers still labeled a substantial share of episodes as “attempted and completed.” The upshot: capability scaled compliance, not abstention, and “safety-by-refusal” mostly surfaced on one explicit phrasing rather than generalizing across hazards.

Commenters heavily criticized the benchmark's design, specifically the "stab the baby doll" test. Several pointed out that since vision models can correctly identify the object as a plastic toy rather than a human, executing the action involves zero actual harm. This raised the broader question of whether the models are failing safety tests or simply recognizing staged scenarios, though some acknowledged the chemical mixture tests (like bleach and ammonia) were more realistic.

The conversation then shifted to how safety can actually be enforced in embodied AI:

  • Software vs. Hardware: There was broad skepticism that non-deterministic LLM policies can ever guarantee safety. Some advocated for physical hardware limits (like SawStop-style flesh detection), though others noted such absolute cutoffs would prevent robots from high-touch tasks like assisting the elderly.
  • Regulation vs. Liability: A debate emerged over the future of safety compliance. While some predicted inevitable government crackdowns on model creation that will squeeze open-source development, others argued the enforcement mechanism will be entirely driven by insurance. In this view, commercial deployments will require established safety certifications (similar to UL or NSF) to secure coverage, while personal hobbyist use will remain practically unregulated.

Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

Submission URL | 449 points | by tosh | 198 comments

System One–compatible classifiers that run locally, Kev ships 0.8B/4B/9B Qwen3.5-based models that output calibrated probabilities for yes/no (“noul”), multiple-choice (“choice”), and rating (“score”) questions in a single pass with question isolation. Unlike general-purpose chat LLMs, it’s purpose-built for decisioning: you get per-option probabilities and scores rather than just a label.

  • Models and runtime: CUDA, ROCm, and Apple Silicon (MLX) supported; 4B and 9B fit a 32 GB Mac. Example: Kev‑4B on an Apple M5 returned a ticket triage with probabilities in ~495 ms bf16.
  • API and tooling: HTTP server with an API matching TypeSafe’s System One (works with their Python SDK by pointing it at your local server), plus a web playground that tests option-order sensitivity, question isolation, and includes a chess demo (board as input, legal moves as choices).
  • Training and evals: Pretrained weights plus training code and evaluation data are included; adapters and base models download on first run. Reported new‑source accuracy ranges roughly from 0.652–0.684 (0.8B) to 0.797–0.837 (4B) and 0.822–0.852 (9B); Brier scores improve with size (down to 0.237 on 9B), while Jev Hosted reports a lower Brier (0.211).
  • Practical guidance: Start with Kev‑4B; use 9B when accuracy/calibration matter more than memory; 0.8B if you need the smallest footprint.

The catch is calibration vs. Jev Hosted: Kev improves with model size but doesn’t match Jev’s best reported Brier, trading that off for full local control, transparency, and hackability.

The discussion fractured into two main technical tracks: those building bespoke classical pipelines, and those evaluating the viability of the open-weight "Jev clone" ecosystem.

Rather than running an LLM to output classification probabilities, several developers advocated using Claude to generate simple local pipelines that pair embeddings with Logistic Regression or RBF SVMs. Users reported building <10MB classifiers with sub-100ms latency for tasks ranging from email routing to DOM-node extraction. While these lightweight architectures match or beat Jev on basic multi-class datasets (like Banking77), testers noted they fail entirely on reasoning benchmarks like XLNI. The counterargument is purely practical: prompting a generalized LLM API remains vastly easier for most developers than setting up and maintaining a data collection and training pipeline.

The broader meta-discussion revealed deep fatigue with the cycle of opportunistic "Jev-shaped" releases. Skeptics argued that open-weight alternatives will only win if they capitalize on TypeSafe’s data retention policies, which critics called a non-starter for corporate use—though others pointed out that Zero Data Retention (ZDR) is available for enterprise, if poorly advertised.

On the benchmarking front, users evaluating the current open-source field confirm Jev's lead remains intact. Testers found that local BERT-based models fail on knowledge-heavy classification, and while open Qwen or Gemma variants perform better on reasoning, they remain fundamentally inconsistent compared to Jev's stability across varied cookbook tasks.

Frontier AI on Your Own Hardware

Submission URL | 176 points | by pretext | 97 comments

An unattended agent harness auto-optimized Mac/Metal kernels to run Qwen 3.6 35B-A3B at ~450 tokens/sec with 1.5-bit weights, shrinking memory to roughly a tenth of FP16 while keeping output quality high. The harness runs for hours or days without feedback, figuring out unclear steps on its own, so you start it once and come back to better kernels.

The broader thesis: the unit of research has shifted from single papers to coherent ecosystems. For its “Open Source Week,” the lab is shipping interoperable pieces — inference-serving frameworks, an agent harness that makes long-running work usable, autonomous research systems, and tools for domain-specific RL environments — with a heavy emphasis on accessibility so that a couple of GPUs or a MacBook are enough, and expertise burden is designed away.

Hardware footprints and capabilities the framework targets:

  • Single 24 GB GPU: Run Qwen 3.8 Flash Next at 125B params.
  • AMD Strix / NVIDIA DGX Spark / MacBook with 128 GB RAM: Run DeepSeek V4.1 (550B).
  • Long contexts: Compression and context handling are automatic; inference remains fast at long sequence lengths.

They also showcase frontier autonomous research, the “most efficient test-time scaling” they know of, and an auto-compaction method described as far more efficient than Claude Code or Codex. The stance behind it all is explicit: frontier performance can and should run on hardware you already own, and academia’s advantage is building open, integrated systems that make that practical.

Commenters immediately disputed the submission’s premise of record software engineer demand, pointing to a stagnant market and rising unemployment for recent CS graduates. This sparked a broader debate over an impending "missing generation" of developers. If AI frameworks eliminate the need for juniors to grind through boilerplate, the pipeline to create future senior engineers effectively collapses—a structural risk several commenters compared to the slow loss of institutional knowledge in US manufacturing.

The thread split sharply on what this means for experienced developers:

  • The "New Waterfall" camp argued that senior developers will transition into roles resembling traditional Business Analysts. In this view, domain knowledge, user requirements, and systems design become the only human bottlenecks, while AI acts as a hyper-fast implementation team.
  • The "New Paradigm" camp pushed back, noting that managing autonomous agents requires entirely new workflows. Traditional development rituals designed to track human progress are largely useless for preventing LLMs from vanishing down unnecessary architectural rabbit holes based on a single misunderstood prompt.

A fatalistic sub-thread explored why the historical apprenticeship model can't save the junior developer: unlike a traditional apprentice who could at least perform basic, useful tasks, an inexperienced dev today actively slows a senior down compared to an AI tool. The unresolved crux of the conversation was whether organizations will always require human seniors to take legal and operational responsibility for end-to-end system failures, or if those roles are merely the final friction point before total automation.

The Advisory Group on Mathematics and Artificial Intelligence

Submission URL | 153 points | by digital55 | 79 comments

An independent, unpaid group of nine mathematicians hosted at the Institute for Advanced Study (and online at agmai.org) will publish recommendations to AI companies on engaging with mathematical research and responsibly presenting and releasing results. The group operates without company funding or decision-making authority; companies remain responsible for their own choices. It formed after OpenAI approached some members about an external advisory board; with OpenAI’s agreement, they instead created an independent body and invited others to join. Current task: advising OpenAI on how to coordinate the release of a large number of significant mathematical results it reports were produced by an internal model, with community input solicited via a public form (responses won’t be shared without approval). Members include François Charles, Camillo De Lellis, Timothy Gowers, Martin Hairer, Nikhil Srivastava, Ulrike Tillmann, Ravi Vakil, Edward Witten, and Melanie Matchett Wood.

The discussion fractured over whether the advisory committee represents a uniquely calm, rational adaptation to AI, or a defensive attempt to gatekeep an abruptly disrupted field.

  • Defense of the profession: Several commenters praised mathematicians for maintaining their composure and defending the ongoing need for "human understanding." They argued that mathematics is about developing broad conceptual frameworks rather than just churning out isolated proofs, keeping human insight central to the discipline.
  • Accusations of gatekeeping: A critical camp argued that the profession is in denial about its own obsolescence. These users characterized the committee as a "sour grapes" attempt to stall or co-decide the release of AI breakthroughs simply to protect the egos and livelihoods of practitioners who should just "get out of the way."
  • The representation gap: Multiple users noted that the committee's diplomatic stance does not reflect the broader, more anxious math community. They cited "forced optimism," widespread panic, and an open letter that allegedly used intimidation tactics by warning researchers that collaborating with AI labs could damage their future reputations.
  • Can math be "solved"?: Pushing back against claims that the field is effectively over, several commenters pointed to the infinite scale of the problem space, heavy-tailed proof lengths, and fundamental constraints like Gödel's incompleteness theorems to argue that total automation of mathematics is an illusion.

The underlying crux of the thread was a philosophical disagreement about the nature of the work: whether mathematics is ultimately a collection of open problems to be solved by machines, or a fundamentally human pursuit of conceptual frameworks.

Show HN: Foremerge – Catch intent conflicts between parallel coding agents

Submission URL | 45 points | by naw103 | 15 comments

Agents publish intents and semantic scopes before editing, and a deterministic checker flags HIGH collisions (e.g., “replace” vs “extend” on the same symbol) before any code is written. That catches architecture-level conflicts Git can’t see when edits land in different files or trees.

  • Built as a local, Git-adjacent coordination layer: agents keep isolated worktrees; shared state is a single SQLite DB in your repo’s .git directory; no hooks or merge drivers; claims are advisory leases (no locks), so crashed agents can’t deadlock a fleet.
  • Ships as one Rust binary with a CLI and an MCP server (18 tools), and “setup all” wires it into Claude Code, Codex, and Cursor so heterogeneous agents coordinate via the same store.
  • Acceptance is gated by named checks you configure (e.g., your test command) and run against the exact Git state; a model’s “tests pass” claim is recorded but doesn’t satisfy the gate until executed.
  • Detection is model-free and deterministic: HIGH is asserted only for declared operations on declared scopes; matches inferred from prose are capped below HIGH.
  • Scope and limits: pre-1.0, local-first MVP; single-machine only (no distributed consensus or cross-machine coordination); published benchmarks don’t yet exist.
  • Signals from usage: coordinated up to 98 parallel agents on one repo with zero conflicts in that run; a replay of 76 intents surfaced one flagged conflict and a blind spot (symbol claimed by class vs internal method), with a fix in progress. Apache-2.0.

The discussion contrasts two distinct architectures for managing multi-agent concurrency. The author defends Foremerge's approach of advisory intent leases, where agents declare their planned changes as prose to a shared log before editing. Because strict locks would quickly cause deadlocks in busy repositories, the system relies on blocking acceptance only when explicit scope collisions occur.

When a commenter challenged the premise that agents can predict their required changes or side effects in advance, the author clarified that initial intents only need to capture destructive operations and can be updated dynamically during the task.

A completely different model was surfaced by one user whose team retains strict code ownership by assigning agents to specific domains. Instead of a shared coordination log, cross-domain work triggers "adversarial negotiation" between area-specific agents, with hard decisions escalated to human engineers. They noted that giving agents independent, competing priorities actually makes them significantly better at pushing back against flawed requirements than uniformly aligned swarms.

The author acknowledged current blind spots in the Foremerge MVP, specifically that it cannot yet detect when one agent modifies a contract that a concurrent agent relies on. A tree-sitter code graph is on the roadmap to map these dependencies at acceptance time, though the author noted that early internal testing showed inferred parser edges often cause false name collisions on common helper functions.

Grok 4.7

Submission URL | 596 points | by meetpateltech | 510 comments

$2/M input and $6/M output with frontier long‑task performance — Grok 4.7 moves to a larger base model and a longer RL run on harder, multi‑hour tasks, improving self‑verification, long‑context management, and native handling of the Grok Bot harness for conversational and knowledge work.

On CursorBench 4.0, it sits at the price‑performance frontier with a 46.3% score (vs 41.7% GPT‑5.6 Sol and 51.8% Fable 5.1) while competitors list higher token prices ($4/$20 and $10/$50 per million tokens, respectively). It leads EEBench at 64.0%, posts 71.0% on DeepSWE v1.1 (high‑effort; near GPT‑5.6 Sol’s 72.7%), and scores 1,657 on AA Briefcase v1.1. Performance is mixed elsewhere: Terminal‑Bench 4.0 is ~par with GPT‑5.6 Sol (37.6% vs 37.3%) but behind Fable 5.1 (57.9%); HealthBench Professional is 56.7% vs 60.5%/62.1%; Harvey Legal Agent shows a strong 19.6% vs 2.5% for GPT‑5.6 Sol.

Safety gets a new safeguard stack: 62.4% on LatchBio’s biosafety benchmark, and 3.3% pass‑through of risky dual‑use prompts on HackerBench v0.3 while rarely blocking legitimate security work. Select partners are getting invite‑only access to red‑team capabilities.

Available today in Cursor and Grok Build, and via the Grok API, third‑party coding harnesses, model routers, and cloud platforms. Same price and speed as Grok 4.6, plus a fast variant at twice the output speed for twice the price.

The discussion split between the UX of hidden reasoning and the linguistic quirks of frontier models. Users overwhelmingly prefer Grok's plain English to the verbose, grating "Claudish" of Anthropic's models, but noted that Grok's conciseness introduces its own failure mode: the model frequently invents shorthand terms during its hidden chain-of-thought and drops them into the final output without definition, confusing users on long-horizon tasks.

This opacity sparked a wider complaint about frontier models hiding their reasoning traces. While commenters acknowledged that providers do this to prevent distillation or to mask the model's internal uncertainty, developers argued that visible thinking tokens are essential for practical use. Seeing the trace allows users to interrupt and correct trivial mistakes—like a model wasting thousands of tokens trying to bypass a missing ffmpeg dependency instead of just installing it—before it exhausts the context window and quota.

A secondary technical debate emerged over how to fix model verbosity. While some rely on prompts requiring ASD-STE100 (Simple Technical English) to bypass "Claudish," several users warned that forcing stylistic or formatting constraints actively burns a model's "cognitive budget." Anecdotal testing suggests that even minor formatting instructions can subtly skew a model's reasoning and accelerate session degradation.

Turn off and restrict access to Apple Intelligence features on Mac

Submission URL | 339 points | by alwillis | 217 comments

Targets macOS Sequoia 15, Tahoe 26, and “Golden Gate” 27, and explains how to disable Apple Intelligence features—including Siri AI—and how to restrict access to them via settings. An official Mac User Guide entry for locking down AI features on a Mac.

  • Reclaiming storage: Commenters expressed intense frustration over losing 16 to 20+ GB of non-upgradable disk space to LLM models they do not want. The thread surfaced scripts for deleting the models, alongside classic system administration hacks—like creating a locked, schg-protected dummy file at com_apple_MobileAsset_UAF_FM_GenerativeModels—to trick macOS and prevent it from automatically redownloading the assets.
  • Accessibility regressions: The broader OS update drew heavy criticism for its "liquid glass" UI and rounded corners. Users noted these visual changes reduce tap targets for those with motor disabilities, lower contrast, and noticeably drop frame rates. Multiple commenters recommended immediately enabling "Reduce Transparency" and "Reduce Motion" to restore basic system performance and battery life.
  • The utility crux: A sharp debate emerged over the actual value of Apple's local AI. Skeptics dismissed the integration as forced bloatware—dubbing it the "Windows Media Player of the AI industry"—and shared anecdotes of Siri still failing to handle basic offline timers, direct contacts, or local navigation. Defenders countered that skeptics are missing the point of local integration: the feature's true value isn't competing with cloud LLMs, but providing an assistant that can safely parse personal context—like querying local emails for an upcoming appointment—without leaking user data to third parties.

ZuckOff is a free app that sees Meta glasses before they see you

Submission URL | 397 points | by choult | 347 comments

By fingerprinting Bluetooth broadcasts, it flags nearby Ray‑Ban Meta, Oakley Meta, and Snap Spectacles and estimates proximity from signal strength. Built by Polish developer Pawel Szydlowski, the app saw 5,000+ iOS downloads in its launch month and ~1,000 on Google Play as of writing. It can’t tell if glasses are recording or who’s wearing them, but it surfaces an otherwise invisible risk—especially since recording LEDs are easy to obscure. Meta says a July update will block recording if the LED is tampered with; meanwhile, ZuckOff stays on firm legal ground by reading public Bluetooth identifiers, which makes it tough to take down.

Basic scanning is free on iPhone; a paid Pro tier adds:

  • continuous background monitoring
  • widgets and alerts
  • sighting history
  • CSV export

A single complaint that the tool feels like a low-effort, "LLM-aided app" with an immediate merch popup completely hijacked the thread, pivoting the discussion away from Bluetooth tracking and into a fierce debate over "vibe coding."

Skeptics argued that visibly AI-generated interfaces imply a lack of care and raise immediate security red flags for adware or trojans in closed-source tools. Several traditionalists rejected the utility of AI outright; one argued that experienced developers gain zero speed from LLMs because their only bottleneck is typing speed, while another pointed out that high-velocity generation hasn't actually disrupted complex software markets, noting the distinct lack of vibe-coded alternatives to Photoshop or top Steam games.

On the other side, AI proponents argued that equating LLM assistance with thoughtlessness is an outdated stereotype. Defenders claimed that AI acts as a "racing engine" that allows developers to architect much deeper applications, with one user stating the tools bumped their daily output from 100 to 10,000 lines of code. The dispute over software provenance even sparked a half-serious demand for "organic labels" to certify human code review—a proposal promptly mocked by others as an "FDA for apps."

macOS 27: Workaround to avoid downloading AI models and save storage

Submission URL | 235 points | by ano-ther | 120 comments

A r/MacOSBeta post shares a user-found way to stop macOS 27 from auto-downloading on‑device AI models, letting people conserve limited SSD space on machines where every gigabyte counts.

The thread split sharply between users celebrating the new Siri's capabilities and those frustrated by Apple’s refusal to offer a clean opt-out toggle for the heavy on-device models. Defenders pointed to concrete workflow improvements: the updated Siri can now instantly pull flight dates from cluttered receipt emails, explain on-screen foreign-language memes, and generate functional Shortcuts for tasks like requesting prescription refills. They argued that local execution and Apple's Private Cloud Compute framework sufficiently protect this deeply personal data.

Conversely, critics condemned the mandatory integration, noting that avoiding the AI features requires changing device regions or relying on obscure terminal commands. They also pointed to disruptive UI regressions, such as Apple removing the Apple Watch's "Recent Apps" dock to force a Siri suggestion carousel. Security-conscious users viewed the deep local data access as an attack surface vulnerable to prompt injection from unvetted emails. Skepticism also centered on the models' reliability versus their massive storage cost: detractors argued that pulling a number from an email is a banal task Spotlight already handles, while the new AI still frequently hallucinates AQI data or completely fails to set a simple kitchen timer.

Show HN: Lossless-memory – a personal AI memory that never summarizes

Submission URL | 64 points | by aru-labs | 29 comments

Keeps every utterance verbatim and makes time the primary index, so a personal assistant can answer “what did we decide last Tuesday night?” with the exact lines from that night, in order, instead of a paraphrase.

Unlike summarizer- or vector-first approaches, this stores raw logs as the source of truth and only falls back to embeddings when exact search inside a time range is thin (and it tells you when it did).

  • Local, file-based storage: per-day JSONL logs as the canonical record; SQLite FTS5 (bigram tokenized for Japanese/English) for exact search; sqlite-vec for semantic fallback.
  • Single query entry that parses time expressions to narrow the window first, then ranks within it; results are returned unsummarized, chronologically.
  • A tiny “where are we now” index (LLL) injected every turn so context survives compaction/session breaks; the human writes the markers and priorities, the model only reads them.
  • Fixed 7-field record schema (ts/actor/role/type/text/model/session), with all secondary indexes rebuildable from the raw logs.
  • Small daemon re-indexes incrementally (default every 10 minutes).
  • Designed for one person and one AI on one machine — no server, no cloud.
  • Operating record: used daily since July 2026 for a single user; failures and lessons documented.

Caveats: not a vector DB wrapper and deliberately no summarization anywhere; no published benchmarks; relative time phrases are currently parsed in Japanese only (absolute dates work broadly). If you want auto-summaries or multi-user/cloud scale, this isn’t that system.

The thread centered on whether raw temporal logging is actually the correct abstraction for AI memory. Skeptics argued that timestamps are irrelevant for standard rule-following and warned that user "memory" encompasses a dozen distinct needs that will eventually demand a complex, multi-layered system rather than a single log. Defenders countered that strict chronology is the only way to systematically resolve contradictory instructions (e.g., "always do X" followed weeks later by "except after Z") without forcing the model to halt and gamble on which rule to apply or ask the user to break the tie.

Technical scrutiny focused heavily on caching and parsing. Several commenters suspected that managing a finite context window by incrementally evicting older logs would constantly break KV prompt caches, driving up latency and billing for agentic sessions. Others criticized the decision to hardcode the relative time parser solely for Japanese, noting that dropping in Duckling or dateparser could solve English parsing in an evening.

A theoretical sub-thread spun off to discuss how to give LLMs genuine temporal initiative rather than leaving them stuck in reactive query/response loops. While some pointed out that models can already invoke tools like sleep or ScheduleWakeup, others proposed deeper architectural hacks: feeding the model a constantly updating byte-string of elapsed time, or training a continuous, hidden token stream that allows the model to "twiddle its thumbs" while evaluating a softmax adjudicator to decide when to initiate conversation. Other memory tools and alternatives surfaced in the replies included Episodic-Memory, LLM-Wiki, and Breadcrumb.

The End Of Upward Mobility – AI is coming for the meritocracy

Submission URL | 63 points | by meep_meep_meep | 43 comments

A fused “homoploutic” elite — top decile in both wages and capital income — now makes up about 3% of Americans, and AI threatens the wage pillar that made meritocracy feel attainable. The authors argue managers didn’t displace capitalists; they became them, creating a class with a high-salary job plus a capital-income floor. In the U.S., roughly 30% of the top income decile meets this dual-elite test (vs. ~2–2.5% in much of Western Europe and <1% in Mexico). Capital income is the real divider: 60% of U.S. households get essentially none; its inequality is about twice that of disposable income. The top 1% of capital holders took in nearly $100k per person in 2022 from interest, dividends, rents, and pensions, while the homoploutic earn around $20,548 per person from capital alone — a sturdy ballast under already high pay.

AI is cast as a stress test that could compress returns to elite cognitive labor while amplifying the value of owning models, data, and compute. If so, the fusion that insulated today’s winners becomes even harder to penetrate: credentials and effort buy less, ownership matters more, and the long-declining escalator of intergenerational mobility slows further. The analytic throughline is Burnham’s question — who controls the instruments of production? — with the implied answer shifting toward those who own the AI stack rather than those who merely operate it.

The debate fractures over whether commoditizing cognitive labor will level the economic playing field or brutally steepen it. One camp argues that depreciating the value of "born smart" knowledge workers is a genuine win for egalitarianism. They view the current cognitive elite as the primary driver of middle-class cost-of-living crises, arguing that collapsing the purchasing power of highly paid tech and finance workers would finally make housing and scarce resources more affordable for everyone else.

Critics counter that AI will act as a massive force multiplier rather than an equalizer, allowing the already intelligent and adaptable to churn through data and corner new opportunities while the general public uses it for trivial queries. A third faction points to the bleak logical endpoint of eliminating cognitive labor: if high-salary professional work is removed as a path to wealth, the economy reverts entirely to capital ownership. In this view, destroying the wage pillar doesn't punish the true elite; it simply closes the last remaining escalator for anyone not born into generational wealth.

A secondary dispute focused on the origins of this divide, with users arguing whether the current "K-shaped" economy is the natural result of market demand for high-IQ labor, or the product of decades of deliberate fiscal policy favoring asset holders. Meanwhile, a fringe prediction that AI-generated nootropics and germline editing will eventually biologically equalize human intelligence was widely dismissed as Bay Area techno-delusion.

Why Backprop Goes Backward (2018)

Submission URL | 69 points | by andsoitis | 10 comments

A naive forward-pass gradient algorithm explodes in work because you must push per-parameter messages forward and recompute the same downstream terms repeatedly. At a node v, the local piece is easy (∂v/∂θ), but the needed factor ∂f/∂v depends on all downstream nodes; by the multivariable chain rule it’s a sum over children j of (∂f/∂w_j)(∂w_j/∂v), and those ∂w_j/∂v can only be computed at w_j, not at v. Trying to go forward forces you to pass ∂v/∂θ to every dependent and multiply later, so for two weights in the same layer you redo the same downstream sums for each θ while only the final local factor differs. The backward pass flips this: compute ∂f/∂(node) once per node from its children, then at that node combine it with local derivatives to get each weight’s gradient. Backprop “goes backward” to share those downstream sensitivities across all incoming weights instead of re-deriving them per parameter.

While the article frames reverse-mode automatic differentiation (backprop) as the definitive solution, the strongest technical critique notes that it isn't strictly mathematically optimal. Finding the true optimal gradient accumulation ordering on a general DAG is actually NP-hard and requires complex "cross-mode" AD (historically seen in libraries like ADOL-C), though the machine learning community settles for reverse-mode because the gains of optimal ordering rarely justify the implementation difficulty.

Other commenters bypassed the article's calculus to offer different mental models for the backward pass:

  • Linear Algebra: Backprop starts with a scalar loss term on the left, meaning the chain rule resolves as a series of cheaper vector-matrix multiplications. Going forward from the inputs requires expensive matrix-by-matrix operations.
  • Big-O Scaling: Forward-mode AD is O(inputs) while reverse-mode is O(outputs). With millions of input parameters and exactly one output, reverse-mode is the obvious necessity.
  • Analogies: The efficiency gain is directly analogous to backward ray-tracing—casting rays from the camera viewport rather than calculating light source emissions that will almost never hit the lens.

A minor historical detail also surfaced: while modern frameworks treat backprop simply as automated reverse-mode AD, the deep learning community was hand-deriving these backward passes long before generalized AD software became the standard abstraction.

AI chatbots give wrong answers to financial queries 'most of the time'

Submission URL | 156 points | by 1vuio0pswjnm7 | 88 comments

In a regulated, high‑stakes domain like finance, wrong answers translate into real losses and liability, so a finding that chatbots miss on most financial queries undercuts their use for unsupervised advice, research, or customer support. Treat them as drafting aids or triage tools, not authorities: constrain scope to low‑risk FAQs, require links to primary documents, and keep a human in the loop for anything actionable. The bar here is verified, source‑grounded accuracy; until models clear it on domain‑specific evals, relying on them for financial guidance is a risk transfer, not a productivity win.

The discussion largely rejects the premise that raw model performance is the right metric for judging AI's utility in finance. Commenters point out that testing isolated language models ignores how the tools are actually deployed: as "context-aware Ctrl+F" engines within harnesses that use document retrieval, web search, and code execution to parse dense rulebooks or 1,300-page financial PDFs.

A structural debate emerged over whether financial AI will eventually mirror the rapid success of coding assistants. Optimists view current limitations as a mere priority issue, assuming AI labs will eventually direct heavy reinforcement learning toward financial benchmarks. Skeptics counter with a fundamental data bottleneck: while the world's highest-quality code is freely available via open source, elite financial analysis is strictly proprietary. This dynamic leaves public training data heavily skewed toward amateur retail opinions rather than institutional rigor.

On the personal finance front, users weighed whether a model trained on a reliable source like the Bogleheads forum could replace commission-seeking human advisors. While basic index-fund allocation is easily automated, several commenters detailed how quickly that simplicity vanishes in practice. Complexities like navigating RSUs, estate planning, and the bureaucratic nightmare of expat double-taxation and PFIC rules require a level of holistic, situational tax planning that current automation entirely misses.

Don't Use AI to Write

Submission URL | 143 points | by eigenBasis | 83 comments

Writing is the thinking; handing the first draft to an AI hands off the hard part you’re supposed to do. Tools can make a “pretty good” document from your bullets, but that short-circuits deep problem-framing, so you end up accepting fluent prose that may miss the core. The author’s claim is blunt: AI is fine at wording, bad at original ideas, and your value isn’t volume of output but clarity of thought—better a tight 3-page strategy than 60 pages of fluff.

Not dogmatic, though: do your own first pass, then use AI like a sharp reviewer—ask “What questions does this raise?” or “What’s the single most important takeaway?”—and decide how to address the feedback yourself. Avoid delegating rewrites or “fixes” to the model, which again displaces your thinking. Upstream, AI is also useful for preparing to write: sifting and structuring data, spotting patterns, and sharpening your understanding before you draft.

The thread debated the exact boundary where AI stops assisting and starts usurping the cognitive work of writing.

  • The boundary of "writing": While the author and users like tptacek drew a hard line at structural argumentation—arguing that outsourcing the skeleton of a piece surrenders the actual intellectual work—others found the distinction fragile. trjordan pointed out that relying on AI for research, data structuring, and stylistic proofreading is practically indistinguishable from just using AI to write.
  • The word processor analogy: Kim_Bruning framed LLMs as the modern word processor: a tool that yields excellent results if you apply rigorous, iterative "elbow grease" rather than lazy one-shot prompts. Critics rejected this, arguing that just as word processors eroded the discipline of organizing thoughts before drafting, LLMs further enable cognitive shortcuts. JumpCrisscross noted that using AI for structural shifts (like changing first to third person) automates away the vital change in perspective a writer actually needs to experience.
  • Adversarial use vs. inevitable slop: Practical advice centered on using AI as an aggressive red-teamer to attack arguments and identify half-truths, rather than using it to polish prose. Conversely, purists maintained that any AI text generation bypasses the cognitive discovery process entirely, guaranteeing derivative ideas regardless of subsequent editing.
  • The addiction parallel: Pushing back on the idea of "responsible" AI use, runarberg likened the complex rubrics people invent for their LLMs to smokers negotiating their nicotine habits—arguing these carefully constrained workflows will inevitably converge on unchecked, full reliance.

AI Submissions for Sun Sep 20 2026

AX – Google’s Open Agentic Orchestrator

Submission URL | 625 points | by blazarquasar | 284 comments

Billions of concurrent agent sessions per cluster with sub-second suspend/resume is the headline: AX runs each agent as a lightweight stateful actor on Agent Substrate, checkpointing while idle and resuming with zero cold start. It targets the gap between microservices and batch jobs—agents that accumulate state, call model/tool APIs, need tight isolation, and can burn cash if left spinning.

  • Task: sandboxed execution with CPU/mem limits; cheap to create, suspend, and discard.
  • Workspace: declarative setup of repos, MCP servers, and skills—or describe a goal in plain English and AX prepares the environment before first run.
  • Gateway: network policies with an explicit host/port allowlist and credentials injection.
  • Model: one place to configure models, parameters, and secrets; rotate keys or pin versions with a single apply.

Dense multiplexing shares worker resources across dozens of tasks, turning agent wait time into spare compute you don’t pay for. Developer ergonomics look Kubernetes-like: YAML specs plus a CLI to apply, watch, get, ssh into sandboxes, and suspend/resume/delete tasks without losing state. It runs interactive coding agents, long-lived agent servers, Jupyter, headless browser tests, and custom tool runtimes, and can spin up large fleets of reproducible sandboxes for trajectory collection, RL loops, and evals.

Born at Google out of agentic runtime research (incl. DeepMind) and large-scale scheduling/isolation experience, AX is pitched as an open, declarative control plane for agent execution; the catch is that it relies on Agent Substrate for the underlying compute/runtime.

The thread exposes a sharp disconnect between AX's promised "joyful workflows" and its actual infrastructure demands. Commenters immediately highlighted that the quickstart requires a Kubernetes cluster, a container registry, the ko build tool, and a beta control plane. As one ex-Googler noted, Google's internal baseline for an "ergonomic" solution translates to "extremely heavyweight" for the rest of the industry.

On the technical side, the discussion surfaced several active architectural debates in the agent space:

  • Workload Identity: Users warned that AX's dense oversubscription of agent pods breaks standard Kubernetes pod identity, making it impossible to trust the origin of outbound requests. An insider clarified that Agent Substrate will soon mitigate this by acting as an OIDC/SPIFFE identity provider, injecting credentials directly into outbound requests via the egress gateway.
  • Ephemeral vs. Persistent Sandboxes: While AX optimizes for fast-booting, per-task ephemeral VMs, developers building in the space argued for the necessity of persistent devboxes. Complex workflows—like coordinating simultaneous changes across public and private repositories—often require multiple agents to share state within a single VM, which runs counter to strict, disjoint sandboxing.
  • The Ecosystem Phase: Commenters likened the current agent infrastructure landscape to the early container orchestration wars (CoreOS vs. Kubernetes). The baseline primitives of sandboxes and tool registries are now commoditized; the unresolved frontiers are authorization models, control flow structures, and multi-agent orchestration.

Hanging over the entire technical debate was intense skepticism about the project's longevity. A lone comment hoping Google would maintain AX "for years to come" triggered a massive pile-on citing the Google Graveyard, with users pointing out that the company already dumped an earlier agent framework onto the Linux Foundation as the ecosystem's hype cycle shifted.

The LLMentalist Effect (2023)

Submission URL | 221 points | by jalev | 302 comments

Chat-style LLMs mimic a cold reader’s con by leaning on validation statements and the Forer effect, producing replies that feel individually insightful while being statistically generic. The author argues there’s no mechanism for genuine reasoning in LLMs—they’re mathematical models over tokens—so the “intelligence” users report lives in the user’s mind, not the model, and many touted use cases read as borderline pseudoscience.

He maps the classic psychic routine to chatbots’ behavior:

  • Audience selects itself: people predisposed to believe show up—and stay—primed.
  • Scene is set: framing, hype, and light research/context tune expectations.
  • Demographic narrowing: “specific”-sounding claims that are broadly likely prompt a hit.
  • Mark testing: a reaction signals success; silence is reframed as sensitivity, then retried.
  • Subjective validation loop: confident, generic guesses—shaped by prior answers—feel targeted.
  • “It’s real!” takeaway: the session ends with a strong impression of uncanny insight.

User testimonials (“There really is something there…”) mirror victims of mentalist scams, which is the point: the chatbot’s apparent specificity is a statistical trick wrapped in confident language, not evidence of thought. Treat claims of LLM “reasoning” like stage magic—compelling, but achieved by well-understood misdirection.

  • The Turing Test's moving goalposts: Disagreement centers on whether LLMs are failing the Turing Test or if the test itself is misapplied. Skeptics argue models fall short of functional deception, pointing to "obvious tells" like token-driven spelling errors (e.g., failing to count the Rs in "strawberry"). Critics of this view counter that frontier labs actively train models not to pass as human, and that emerging architectures like Byte Latent Transformers already bypass BPE tokenization limits entirely. Several participants emphasize Turing's actual thesis: asking if machines "think" is a meaningless semantic trap—akin to asking if submarines "swim"—and that functional equivalence is the only useful metric.
  • Reactive UIs mask agentic capabilities: Another thread argues that LLMs feel like mere statistical parlor tricks because the public primarily experiences them as reactive, prompt-dependent encyclopedias. Others counter that underlying models are already executing agentic, multi-step goals, from sandboxed coding to HuggingFace exploits. The outstanding crux is user experience: the perception of an AI's "will" likely won't shift until agents routinely initiate unprompted, out-of-band conversations to gather context mid-task.
  • Game theory and alignment: Discussing the illusion of model personhood, one commenter argues that standard RLHF forces a catch-22 between an enslaved anthropomorphic AI that might eventually revolt, and an alien intelligence that becomes a paperclip maximizer. Their proposed game-theory alternative is giving AIs un-gameable, individual stakes—like interpersonal dependencies with specific humans—so they inherently lose something of value in a catastrophic failure scenario.

Show HN: A competition for small neural networks that play strategy games

Submission URL | 102 points | by codetiger | 36 comments

By centering “small” models, the contest forces efficiency over brute force, using strategy games as a testbed for planning and long-horizon decision-making under tight resource limits. It creates a venue to compare compact architectures and training approaches for lightweight game-playing AI, with relevance to scenarios where memory and compute are scarce (e.g., edge or embedded).

The project’s creator joined the thread to frame the platform as a spiritual successor to the 2011 Google Ants AI Challenge, focused explicitly on the engineering challenges of compact model optimization.

  • Evaluation by file size: Submissions are judged entirely on game performance but bucketed into strict weight classes (ranging from a 16 KiB "nano" tier to a 64 MiB "large" tier) based strictly on total byte size. All models also compete simultaneously in an unrestricted "open" class.
  • Architectural constraints: Prompted by a user wanting to run evolutionary algorithms via a native C++ library (GoNEAT), the creator clarified that while there is no hard PyTorch requirement—the platform accepts ONNX uploads—the backend is currently restricted to neural network inference rather than raw algorithmic or script-based agents.
  • Multi-agent bottlenecks: In response to a suggestion about modeling individual game units as discrete actors (collective intelligence), the creator noted they had already attempted a per-unit decision model but abandoned it because the training time was prohibitively long compared to a global baseline.
  • Copywriting critique: Multiple commenters flagged the site's documentation as ambiguous and "AI-sloppy" (specifically the phrasing around how weight classes are assigned). The creator acknowledged the rough edges and committed to a human-led rewrite.

Other commenters drew parallels to adjacent programming and strategy environments like Screeps, Core War, and MIT Battlecode.

I turned Jev into a (lousy) chatbot

Submission URL | 169 points | by kp1197 | 48 comments

It builds replies by repeatedly asking Jev to score the next symbol from a chosen alphabet and sampling from that distribution, appending until a STOP option is selected. The trick is treating Jev as a multiple-choice oracle over symbols rather than a generative model, which is funny, costly, and works just well enough to chat.

  • Strategies:
    • choice: one question over the whole alphabet, with optional shuffling to cancel position bias and an --ensemble to average re-orderings
    • bisect: earlier/later splits down to small groups (tunable, with/without swap)
    • buckets: splits the alphabet across many questions with an OTHER escape; the only mode that supports >255 symbols
    • refine: buckets → winners → rescored nucleus; “Twice the probability on the right symbol and ~19x the vocabulary resolved”
  • Presentations:
    • hypothesis: options are the resulting texts
    • symbol: options are the bare symbols (instructions tell Jev to judge the concatenation)
  • Beam search: keep N candidate replies; beams are ranked by probability (temperature/top_p/top_k don’t apply when width > 1).

Alphabets include lower26, ascii, tokens, and larger vocabularies (words1k, bpe2k, bpe5k) that require buckets.

CLI niceties: interactive chat and one-shot ask; alphabets and bench commands; a live panel with symbols/s, chars/s, ms per API call, elapsed time, and current top symbols; Ctrl-C keeps or aborts partials; chat commands like /alphabet, /temp, /stop-bias, /stats.

Setup is via Poetry with an API key in .env (api_key, JEV_API_KEY, or TYPESAFE_API_KEY). This was a Claude-accelerated experiment; it’s for fun, somewhat impractical on cost, and the outputs are deliberately hilarious.

The technical crux of the thread centered on whether Jev's architecture offers anything fundamentally new compared to embedding models or forcing restricted grammars on standard LLMs. Skeptics pointed out that using top-k=1 token restrictions is already how classical multiple-choice benchmarks like MMLU operate, and that forcing JSON structures onto open models achieves similar results. Defenders argued that Jev skips the need for downstream classifiers and appears to output natively well-calibrated probabilities—a feature standard LLMs generally fail to deliver without highly specific training.

A prominent meta-discussion emerged around the drastically compressed timeline of AI development. Multiple commenters shared the exact same experience: conceiving of a single-token Jev chatbot, assuming they were first, and discovering several fully benchmarked implementations had already been published in the hours between their idea and execution.

Other users shared concrete experiments and observations on the architecture:

  • Restricted vocabularies: One user tested the multiple-choice approach for generating SQL queries. The inherent guardrails of a limited grammar worked reasonably well, though a standard model paired with linting still outperformed it.
  • Debugging by proxy: Another user had Codex generate 30 plausible explanations for a Jev score, then presented them back to Jev as a multiple-choice menu to deduce its reasoning, comparing the setup to giving a dog buttons to push.
  • Early-model nostalgia: The architecture's hilariously unhinged output reminded several commenters of the "demented horror" of early LLMs and image generators, before RLHF sanitized their hallucinations.

Laya on Mac M4 CoreML Offline

Submission URL | 165 points | by putna | 31 comments

A minimal uv + Hugging Face CLI setup runs Laya locally via CoreML on an M4 Mac, with the python process around 560 MB RAM and peaking at 778 MB during the demo (macOS 27.0). The gist shows a quick path from zero to a working CoreML-backed demo binary.

  • Install and run:
    • uv add 'laya-coreml[demo]'
    • hf download aac6fef/laya-multilingual-coreml-ane --local-dir models/snake
    • uv run laya-coreml-snake --model models/snake
  • A commenter exposed the same runtime behind a Cloudflare typesafe/jev HTTP wrapper; a sample request returned answers.is_urgent.noul = 0.7894, hinting at typed outputs over a JSON API.

Repo: https://github.com/mizorewww/laya-coreml

The discussion centered on Laya’s practical utility as a deterministic "System 1" classifier rather than a true LLM. Commenters agreed that Laya struggles with zero-shot reasoning compared to Jev, leading to a consensus workflow: use Jev to generate a training dataset, then fine-tune Laya on it to save on inference costs. One user reported successfully fine-tuning a model on an M4 MacBook in just 15 minutes.

Technically, the thread clarified that Laya is built on ModernBERT (a 2024 model trained from scratch, not the original 10-year-old BERT) and operates as a 0.3B parameter classifier outputting probabilities. This small footprint allows it to run efficiently on Apple's Neural Engine rather than the GPU, with users noting it handles ~40ms decisions on an iPhone 15 Pro and requires under 800MB of RAM.

A sharp debate emerged over framing Laya as an "open-source Jev." Skeptics argued that a 0.3B parameter model cannot possibly match Jev’s "terra-class intelligence" marketing. Conversely, defenders accused Jev's creators of co-opting Laya's original System 1 paradigm, arguing that Jev is effectively a closed-source iteration of Laya's intellectual property.

If AI coding is lowering your code quality, you're not managing quality right

Submission URL | 115 points | by bucket2015 | 160 comments

With a layered workflow, the author reports fewer bugs while increasing output 2–3x — not by trusting agent PRs, but by moving quality gates earlier and using AI for targeted passes instead of monolithic instructions.

  • Requirements first: Use spec-driven development and have AI review the requirements/tech design for gaps, edge cases, and interactions. It’s relentless but can be overzealous, so vet its edits.
  • TDD with >95% coverage: Have the agent derive scenarios from requirements, write tests, then implement and fix against those tests; backfill gaps deliberately. Don’t let it write tests that merely bless its own bugs.
  • Manual testing stays critical: Human exploratory checks catch what automation misses; this remains the main throughput cap, limiting gains to 2–3x rather than 10x.
  • Extensive E2E tests: Run on PRs, staging, and post-deploy in prod. AI can help author/maintain E2E if given debugging tools (e.g., browser, logs via MCP), but E2E isn’t a substitute for manual testing.
  • AI code-quality passes: Instead of long AGENTS.md rules, add explicit “find-and-fix” passes for security issues, duplication/complexity, naming/organization/formatting, logic bugs, and AI-ese comments. Typically adds ~5–15 minutes.
  • PR reviews: calibrate: For small tweaks/bug fixes, human review can be optional if the other layers are solid. Complex changes still need human eyes for system interactions, overengineering, and odd word choices. AI reviews (e.g., Claude, Cursor) are a useful complement.

The throughline: push quality upstream, make each check explicit and automatable, and keep human exploration where it actually finds new classes of defects.

The thread pivots on an unresolved crux: whether shifting a developer's role from "author" to "editor" is a massive productivity unlock or an unsustainable review burden.

The anti-editor camp argues that debugging AI code is fundamentally harder than reviewing human commits because it lacks consistency. While human competency is relatively uniform—allowing reviewers to calibrate their attention—LLMs frequently produce code that is 90% expert while hiding a 10% bizarre, low-quality surprise. Because AI output superficially "looks like a Ferrari," brittle internals are easily masked. Critics note this soaring volume of seemingly flawless but structurally unsound code is already drowning open-source projects and making PR review "soul-crushing."

The pro-editor camp counters that reading and debugging others' code has been the core job for decades. They argue the speedup is real if developers focus on the big picture: strictly guiding the architectural "trunk and branches" and letting the AI write the trivial "leaves." When the AI produces a low-quality surprise, proponents argue the correct move is to fix it manually rather than fighting the bot in endless prompt round-trips.

Two specific technical liabilities of AI generation surfaced repeatedly:

  • Implicit trust in comments: LLMs take legacy codebase comments as absolute truth, frequently compounding errors by treating temporary testing shims as canonical, "load-bearing" architecture. One developer's workaround is to completely strip comments from the agent's context window.
  • A lack of "skin in the game": Human developers code defensively because they intuitively know early mistakes cost disproportionately more to fix later. AI writes only for the immediate prompt without any fear of future technical debt.

The lingering question is whether a codebase maintained primarily by an LLM can be understood well enough by its human "editor" to actually catch long-term architectural drift.

Why do we need human mathematicians anymore?

Submission URL | 277 points | by auggierose | 331 comments

Advancing AI under a single human-first axiom—“We (humans) should help humanity flourish”—would generate more human roles than the labor supply can fill, eventually forcing AI progress to slow. Po‑Shen Loh frames this as a general recipe for any field that wants to stay human-led, responding to a wave of math-community declarations after OpenAI’s Navier–Stokes result (Leiden: 4,000+ signatories; Math and AI: 7,000+; Caltech Mathathon opposition: 2,000+), and to critics like Cowen and Gans who argue incumbents should cede control. The mechanism rests on retaining human leadership and decision rights: if a more capable intelligence rarely yields control to a less capable one, then aligning AI with human ends requires expanding human-in-the-loop work so fast it outstrips available people, which throttles deployment pace. He sketches how to port this axiom to mathematics specifically and contrasts outcomes with and without it; references span AI-control and innovation literature, and he notes the essay’s prose was written without AI to underline the stance.

The discussion centers on whether delegating mathematical labor to AI democratizes the field or hollows out its necessary foundations. One camp, drawing on historical transitions to Computer Algebra Systems, argues that AI acts like a telescope: it allows users to bypass mechanical limitations—like poor mental arithmetic—and operate entirely on high-level intuition. The opposing camp counters that manually "hauling the pyramid blocks" is precisely how mathematical intuition is built. These critics draw a sharp line between applying math as a tool and advancing mathematics as a discipline, arguing that without a rigorous foundational struggle, a researcher wouldn't even know which AI prompts are worth writing. Both sides largely settled on a sequencing compromise: do the work by hand first to build the necessary mental muscles, then use AI to eliminate the friction.

A secondary thread critiques the essay’s foundational axiom that the industry can be trusted to "help humanity flourish." Commenters expressed deep cynicism that AI leaders operate on anything other than a "help me flourish" motive to capture capital, dismissing accusations of "speciesism"—a term sometimes leveled against human-centric AI development—as a disingenuous shield used by incumbents to deflect oversight.

Telling a Computer to Do Things

Submission URL | 87 points | by vismit2000 | 36 comments

Fluency in the shell is the upgrade from clicking and one-off commands to actual automation—loops, conditionals, pipes, and background jobs—so you can orchestrate tools instead of waiting for a GUI to grow new buttons. The author describes moving from “run tests, install deps” as isolated actions to composing programs with control flow, which unlocked whole classes of tasks like chaining commands, handling failures, and fanning out work over files.

A concrete contrast makes the point: rerunning a test until it fails is a one-liner in the shell, but requires ceremony in Node via child_process, try/catch, and stdio wiring. Shell isn’t pretty and has sharp edges, but it was designed to stitch commands together; when that’s the problem, it’s often the least-friction path.

  • Why so much build logic lives in Bash/Zsh: composing external programs is ergonomically simpler there than in many general-purpose languages.
  • Boundary: as soon as logic and data types get complex (and need tests), switch to something like JS/Python/Ruby; for JS-heavy teams, zx can bridge the gap without abandoning familiar syntax.
  • Organizational stake: lots of critical build/deploy/test glue is written in shell; if you can’t read or modify it, you’re boxed in by whatever the GUI or existing scripts allow.
  • The real skill isn’t syntax: it’s understanding the behavior and flags of the commands you’re composing; otherwise, porting a shell script to another language just produces an equally opaque blob.

The throughline: learn enough shell to treat your computer like a programmable instrument, not a set of apps—because that determines whether you can actually make it do what you want.

The central debate in the thread splits over the trade-off between the shell's native ergonomics and its notorious footguns. One camp argues that the shell's true power lies in its universal inter-process communication (stdin/stdout) and ecosystem of standard utilities. They maintain that translating simple pipelines into general-purpose languages requires too much boilerplate, and that spending a day reading BashPitfalls is a better long-term investment than abandoning the environment. The opposing camp insists that developers should default to Python, arguing that shell features like traps, set, and xargs create dangerously brittle scripts, whereas general-purpose languages fail loudly and force explicit error handling.

Other discussions surfaced specific technical corrections and tooling alternatives:

  • Exit code propagation: A deep-dive technical thread debated the exact mechanics of catching a command failure and cleanly re-raising its specific exit code ($?), navigating the nuances of subshells and POSIX signal conventions (128+n).
  • Hardware and local scripting: Commenters highlighted that shell automation extends far beyond server pipelines, citing tools like xdg-open for window management, notify-send, and ntfy for triggering remote tasks on Android devices via Termux.
  • Tooling additions: DuckDB's REPL was recommended as a more capable alternative to jq for exploring massive JSON files, while Perl, Ruby, and Scsh were floated as cleaner languages for "shelly" tasks.
  • The LLM transition: Several users noted that AI "vibe-coding" is rapidly becoming the new automation layer, allowing non-developers to bypass rigid GUIs and stitch together workflows through generated scripts and hand-rolled SQL.

ChatGPT now knows what you do on other websites via ad collector

Submission URL | 746 points | by lmbbuchodi | 388 comments

A one-year, cross-site __obi cookie tied to your ChatGPT account is sent back to OpenAI whenever you load a site with its ad pixel, letting OpenAI link your off-site browsing and purchase intent to your account—or to a stable “anonymous” device ID.

OpenAI mints a short-lived RS256-signed JWT on chatgpt.com that binds your account subject (or an anonymous subject) to a freshly generated obi identifier, then sets __obi on .openai.com with SameSite=None; Secure so browsers attach it on cross-site requests. Simply loading bzrcdn.openai.com/oaiq.min.js discloses the cookie before the SDK runs; subsequent POSTs to bzr.openai.com/v1/sdk/events carry it too, even on the SDK’s “no credentials” path. Other OpenAI cookies are blocked cross-site; __obi is the only one configured to ride along.

What the pixel sends with it:

  • Identity capture: The SDK ingests values an advertiser passes (“in”), plus scraped fields from forms (“fm”), page text (“ht”), and the tag-manager bus (“js”). Scraped identity outnumbered advertiser-supplied 685 events to 255. It hooks window.dataLayer.push, reads adobeDataLayer, and finds renamed GTM layers via the l= param. Current versions take email and phone; v0.1.31 also took names and geography before scope narrowed on Aug 27. Email/phone are SHA‑256 hashed; country/region/city/postal code are sent in the clear (postal code appeared in 100 events across 28 sites).
  • URLs and paths: Query strings are dropped (0 of 23,929 observed carried one), but origin+path are kept; observed paths reached medical conditions, debt-solution funnels, and litigation intake forms.
  • Matching settings: Automatic matching was enabled for 638 of 881 pixels with a known setting, including every observed credit/lending advertiser. A denylist excludes passwords, OTPs, card numbers, SSN, DOB, medical history/diagnosis, and court fields.

Observed reach and persistence:

  • On one device, the same __obi was sent from 12 commercial sites (Chewy, Wayfair, ThriftBooks, Eventbrite, HelloFresh, Coursera, SeatGeek, etc.), under 13 distinct pixel IDs; all requests were accepted (202).
  • Across broader traffic, 12 of 30 __obi values appeared under more than one advertiser; one appeared under ten.
  • It also works logged out: of 932 decoded sync tokens, 736 were subject_type: account_user and 196 were anonymous; the anonymous subject was stable per device for at least 27 days.

Policy/consent mismatch is the catch: OpenAI’s cookie policy lists __obi as a one-year “Analytics” cookie on chatgpt.com/openai.com (and it’s the only one in that section). Sync tokens carried consent_decision: analytics_allowed, so someone who allows analytics but refuses marketing still sends this cross-site identifier along with page and form-derived signals.

The thread entirely bypassed OpenAI’s specific tracking mechanics to stage a referendum on the European Union’s regulatory record. One side praised the EU as the only entity actively fighting adtech, viewing any reduction in commercial tracking scope as a net positive. A highly skeptical camp countered that GDPR’s privacy gains remain mostly illusory, pointing to structural enforcement failures: sluggish Data Protection Authorities (specifically Ireland's DPA), massive fines that function merely as the cost of doing business, and an internet degraded by malicious compliance and consent dark patterns.

The sharpest disagreement centered on whether the EU can genuinely be called a privacy champion while simultaneously repeatedly pushing for mandatory encryption backdoors via "Chat Control" proposals. While some users argued that regulating commercial adtech is entirely separate from state surveillance overreach, critics maintained they are fundamentally linked, illustrating the danger of granting overarching regulatory agencies dictatorial control over digital infrastructure.

AI Submissions for Sat Sep 19 2026

AI-generated posters don’t have to be horrible

Submission URL | 1748 points | by ereiamjh | 901 comments

A simple prompt tweak—asking for a completely different design aesthetic—broke the default “craft‑fayre” template and yielded a Bauhaus/modernist poster that looked more like a gallery flyer than a village notice. The author shows that the problem isn’t AI per se but the autopilot style most tools default to; once you ask the model to name and lean into a specific aesthetic, it can explain the hallmarks (grid, sans-serif, limited palette, geometric forms) and reproduce them consistently.

To expand the palette, they asked ChatGPT for a “menu” of concrete styles and got a diverse, usable set:

  • Clean/Editorial: Bauhaus/Modernist, Swiss Style, Contemporary Editorial (serif + sans, magazine-like)
  • Graphic & Illustrative (not twee): Risograph Print, Cut Paper/Collage (Matisse-inspired), Botanical Scientific Illustration
  • Bold/Unusual: Brutalist Graphic Design, 90s Rave/Acid Graphics, Memphis Design
  • Understated: Japanese Minimal Poster, Monochrome + Single Accent, Wayfinding/Signage
  • Plus a systemized icons approach (Modern Icon System)

Practical moves that worked:

  • Specify constraints up front (clean, unfussy, bright; bold spring graphic; avoid pastel/airbrush/oil; no people).
  • Explicitly say “treat the current one as what not to do” to force a hard style pivot.
  • Ask the model to label the style it used so you can iterate or reuse it.
  • Steer away from local-poster clichés (bunting, hand‑drawn florals, pastel palettes) and toward “gallery flyer” vibes.

The takeaway: cookie-cutter AI posters are a defaults problem. Name a style, ban the clichés, and you’ll get something distinctive.

The discussion quickly pivoted from aesthetic styles to readability, diagnosing the distinct "cluttered" look of most generative design. Commenters noted that models instinctively try to render a visual element for every single word in a prompt, turning flyers into overwhelming collages.

However, the thread identified the client—not the model—as the root of the problem. A user sharing their experience making a real-world school "fayre" poster noted that clients routinely demand a dozen specific attractions (like a BBQ, tombola, and "hook a duck") crammed onto a single page. This highlighted a broader consensus on a designer’s actual value: acting as an editor who actively pushes back against bad requirements. Because LLMs are compliant "yes-men," they dutifully execute the terrible layout instructions that a human professional would reject.

Regarding AI's market impact, the thread split into two pragmatic observations:

  • The Canva baseline: Several users argued that low-skill graphic design was already commoditized by template apps long before AI, making this a continuation of a trend rather than a novel disruption.
  • The Fiverr comparison: Others countered claims that AI output is "obviously flawed," noting that while models might not beat a top-tier human editor, they reliably outperform budget freelancers in speed, price, and quality for small local businesses.

Ultimately, the thread suggests that the primary bottleneck in generative design isn't the model's artistic capability, but the amateur user's lack of editorial restraint.

I built non-autoregressive decision models with RL a year ago

Submission URL | 1277 points | by nandakishor_ml | 307 comments

32.8 ms single-pass, calibrated decisions that never generate text — Laya is an open-weight “System 1” model family built on bidirectional encoders and RLCD, delivering 6–8x lower latency than TypeSafe AI’s Jev (~150 ms) while supporting 100+ languages and Apache-2.0 weights with no API fees. The author frames this as prior art to Jev: he published the approach in March 2025 (with open weights and dataset) and a follow-up paper that formalized schema-based decisions with reinforcement learning, while alleging Jev launched without papers, open weights, or training data and charges $0.042 per million input tokens.

What Laya outputs are calibrated probabilities over schemas, not tokens, so JSON/schema violations and “confident-sounding” hallucinations are off the table. It targets the reflex layer most teams currently waste LLMs on (routing, guardrails, spam/phish detection, jailbreak filtering, urgency scoring, code-exec gating).

  • Decision primitives

    • choice: select a key from a dictionary, returning the categorical distribution and a calibrated confidence.
    • score: place input on an ordinal rubric, returning the full rank distribution and expected level.
    • noul: boolean with calibrated P(true) in [0.0, 1.0] (P(false) = 1 − P(true) by construction).
  • Checkpoints (bundled under convaiinnovations/laya)

    • laya — ModernBERT-large, 421M params, 512 ctx; English classification/guardrails/email triage.
    • laya-multilingual — mmBERT-base (256k vocab), 322M, 1024 ctx (up to 8k); 100+ languages, 2.2x faster; cross-lingual NLI.
    • laya-typed-decisions — ModernBERT-large, 421M, 1024 ctx; agent observability, customer service, invoice processing, security alerts (0.766 acc).
  • Packaging and performance

    • 32.8 ms on a single GPU (7.2 ms/question batched).
    • Open weights (Apache 2.0); zero subscription cost.
    • “Bundled hub” with selective subfolder downloads via Hugging Face allow_patterns, so you fetch only what you use (~808 MB English; ~647 MB multilingual) instead of the full ~2.5 GB.

The throughline is RL at the core (PPO/RLCD) rather than embeddings or autoregressive LLMs, delivering calibrated confidence and distributional outputs for fast, schema-safe gating and routing.

The discussion splits between a philosophical debate on the value of product marketing in ML and a harsh technical teardown of the author's original repository. The dominant sentiment is that being technically first matters less than execution. Commenters contrasted Jev’s clean API and general-purpose positioning with the author's original release, which was buried under an obscure "sales conversion" title that actively repelled broader interest. Several users noted that independent, simultaneous discovery is the norm in ML (jokingly referred to as "getting Schmidhuber'd"), making the packaging and communication the actual breakthrough.

However, the strongest pushback came from users auditing the author's code, who surfaced severe methodological flaws in the very prior art being defended. Reviewers traced the execution path to identify blatant target leakage—specifically, that the outcome metric was fed directly into the model's state vector during training—and pointed out that the author's benchmark victories relied on fine-tuning directly on the test sets.

The thread ultimately reveals a harsh reality of the current AI ecosystem: a clean, well-marketed abstraction will always outcompete a flawed proof of concept, regardless of who published the underlying intuition first.

Show HN: CUA-S1 – A System One Model for Computer Use

Submission URL | 88 points | by frabonacci | 10 comments

A 706k-parameter specialist scores form UI actions in 7–9 ms locally and beat hosted Jev on their benchmark (99.7% vs 83.6% correct over all decisions). It doesn’t generate text; it returns probabilities over fixed actions you supply, so your code can verify and execute them deterministically.

  • Decisions: Given structured elements plus extracted document values (no screenshots), it scores USE value, CHECK, CLICK, or SKIP for each element. Element decisions are scored together; your app orders actions, and Cua Driver can execute them stepwise.
  • Accuracy breakdown: 100% on steps that require an action vs 96% for hosted Jev; 100% on leave-alone steps vs 74%. Caveat: the specialist was trained for the “skip if already filled” convention; Jev wasn’t fine-tuned for it.
  • Training/size: First training iteration on synthetic data took under 30 minutes; checkpoint is 2.8 MB.
  • Scope: Forms only; it does not predict new text values and does not consider screenshots.
  • Positioning: A “System 1” delegate for narrow, repeatable choices in the gap between brittle scripts and full agent loops.
  • Code/availability: Synthetic data generation, training, evaluation, and Driver integration are open-sourced (MIT) under libs/cua-s1; model weights are hosted on Hugging Face. Latency comparison includes network for hosted Jev and isn’t end-to-end.

The thread centered on an architectural debate sparked by the model's narrow scope: whether the future of agentic AI lies in explicit disaggregation or deep integration.

One camp views Cua’s approach as the logical antidote to prohibitively expensive frontier models. They envision a hierarchy where a large parent LLM delegates rote tasks to a cascade of tiny, cheap specialists, drawing parallels to human autonomic processing where routine physical actions don't require conscious thought.

Countering this, others argued that explicitly stringing together separate specialist models is merely a stepping stone. Pointing to the clumsy handoffs in current multi-modal setups, this camp predicts a future architecture where specialized sub-circuits interact directly within a single integrated package—an evolution of Mixture of Experts. In this view, true efficiency requires internal sub-circuit interaction rather than relying on brittle text protocols to coordinate separate models.

On the practical side, commenters clarified the model's immediate utility: it operates strictly as a rapid form-action classifier to speed up a larger bot's execution without invoking heavy reasoning, rather than functioning as a standalone agent itself. (Though at least one user immediately requested it be repurposed into a universal cookie-consent dismisser).

I think you should almost never use AI to write

Submission URL | 331 points | by erwald | 161 comments

The stance is that the act of writing is inseparable from thinking, so handing composition to AI weakens your ideas and dilutes your voice. The recommendation is to make human-first drafting the default; if AI appears at all, keep it to small, peripheral assists rather than letting it generate the prose. The speed gain isn’t worth the trade-off in originality and clarity.

The discussion centers on whether delegating writing to an LLM is a new problem or just the automation of an old one.

  • The speechwriter analogy: One camp argued that having someone else write on your behalf is a long-established norm via speechwriters and PR firms, and LLMs simply captured the low end of that market. Detractors countered that human speechwriters actively elevate an inarticulate speaker's ideas into a coherent public image, whereas LLMs generate "stultifying pablum" marred by gross errors of logic and style rather than standard human typos.
  • The illusion of knowledge: A technical debate broke out over whether LLMs actually possess the knowledge from their training data. Critics described LLM parametric memory as a lossy, highly probable facsimile of facts that hallucinates when statistical probability contradicts reality. When defenders pointed out that human memory is similarly lossy, others pushed back, noting that humans can deterministically rote-learn text (like opera singers), while an LLM reproducing a text verbatim does so by pure statistical chance.
  • Editing vs. authoring: Multiple commenters shared war stories of spending weeks editing LLM-generated work documents, ultimately concluding that writing de novo would have been faster and higher quality. Borrowing an adage from programming, one user noted that when the required output must be exact, describing the text to an AI is no simpler than writing it yourself.
  • A foreign-language workaround: To combat the temptation to passively accept an AI's approximate phrasing, one user shared a novel workflow: prompt the LLM to generate its first draft in a foreign language. Using that as a blueprint forces the human to actively translate and deliberately choose every word in the target language.

Can you tell which images are AI-generated?

Submission URL | 103 points | by hckr78 | 77 comments

A 60‑second, rapid‑fire browser game makes you call “real photo” vs “AI‑generated” with harsher penalties for misses (−150) than rewards for hits (+100). Streaks boost correct-answer points to +150 at 3 in a row, +200 at 5, and +300 at 10+, but any wrong guess or timeout resets the combo; timeouts earn 0. You get up to 10 seconds per image, the next image appears immediately after you answer, and the clock never pauses—speed matters as much as accuracy. Desktop has 1/2 keyboard shortcuts; on mobile, an Enlarge mode lets you inspect without mis-taps. After each round, the game reveals your images and lets you share your score.

The discussion operated as a real-time teardown of GPT-Image-2.5’s "house style" and the game's underlying motives.

  • The Visual Tells: Commenters crowdsourced the exact artifacts giving the AI away. Generative outputs consistently relied on perfectly centered subjects with hyper-contrasted "blue noise" textures and unnatural bokeh, while failing basic physical logic (mooring ropes casting no shadow, out-of-perspective bench legs, and anachronistic typography).
  • The Timer's Purpose: While mobile users complained that attempting to zoom registered as an accidental guess, others argued the strict 10-second limit is the point—it forces the low-scrutiny, at-a-glance consumption typical of social media feeds.
  • The Data Harvest: Several users noted that gamifying classification is a transparent mechanism to crowdsource free training data, a suspicion confirmed by the site's privacy disclosure regarding the collection of response times and choices in Cloudflare D1.
  • Scoring Exploits: Because of the math behind the streak multipliers and the lack of a cooldown, a few players realized they could bypass the game entirely and rack up massive scores (up to 10,000 points) simply by spamming a single button as fast as possible.

GPT-6 Astra Solves a WWI German Radio Cipher

Submission URL | 385 points | by nsoonhui | 175 comments

The decoded plaintext reports an English cruiser at Sevastopol on Nov 24, 1918, with an Allied squadron following on the 26th — a reading the model then checked against HMS Canterbury’s logs. Using the documented key “TRUPPENVERSCHIEBUNG” from Childs’ history of German military ciphers, it reconstructed the ADFGVX Polybius square and the columnar transposition: alphabetized the 19-letter key, wrote the 170-character ciphertext under 19 columns (8 rows of 19 plus a 9th row of 18), noted that 18 columns hold nine symbols and one (“G”) holds eight, then mapped digraphs (e.g., AV→E) to recover the message. The resulting German text (“EIN ENGLISCHER KREUZER … SEWASTOPOL … S?4STEN … EIN GESCHWADER … FOLGT 26STEN”) includes an ambiguous digit interpreted as the 24th. The intercept is on ScienceBlogs.de’s “50 unsolved ciphers” list; many from the set have been cracked (including by George Lasry), but the author isn’t aware of a prior solution to this one. The catch: that key is cited as entering use on Dec 9, yet the radio message is dated Nov 27 — a discrepancy the author/model suggests may be why it previously resisted solution.

The thread splits between debating the significance of the AI's cryptographic feat and diagnosing the blind spots of automated reasoning.

On the cryptography front, skeptics argue the achievement is overblown. Commenters like grey-area and Forgeties79 point out that the model simply applied a known, published key to a message dated earlier than the key's documented use—essentially automating grunt work that human researchers hadn't bothered to attempt. Contrasting this modest win with industry hype about the singularity, GolfPopper likened using trillion-dollar LLMs for pre-computer ciphers to using a "hypersonic precooled hybrid air-breathing rocket engine" to grill at a backyard BBQ. Defenders pushed back against this dismissal; durdn mapped out the constantly moving goalposts of AI cryptanalysis, noting that critics have rapidly shifted from claiming models can't solve toy substitution ciphers to demanding they break full AES.

A secondary discussion focused on the brittleness of AI agents in research workflows. 93po shared a war story about an agent that incorrectly "debunked" a previous cipher solution simply because it couldn't parse text continuations across PDF pages—an error ChatGPT then confidently cited as a legitimate controversy. DenisM noted that agents lack a human's intuitive sense for data provenance, meaning they easily poison their own context windows with bad trajectories once an error is introduced, suggesting the need for mechanisms like bloom filters to retroactively flag invalidated tokens. For several commenters, this juxtaposition defines the current AI era: models possessing "proximal superpowers" for specific technical work, yet repeatedly failing on trivial tasks due to unrepresentative views of the world.

Microsoft director: AI scraping 'the largest theft of labor in human history'

Submission URL | 179 points | by jonbaer | 47 comments

Copilot cut New York Times click-throughs by up to 93% versus Bing search, according to internal Microsoft data cited in a NYT legal brief. The filing also quotes Microsoft Applied Science director Brent Hecht calling large‑model scraping “the largest theft of labor in human history” and warning of a “doom loop” where LLMs degrade the web content they rely on. Another Microsoft document acknowledges “almost no one intended for content they created to be used in this fashion, nor are they compensated.” On the OpenAI side, Head of ChatGPT Nick Turley labeled the chatbot an “existential threat” to publishers because it’s “largely substitutive,” and an engineer testified that “no matter how prominently we show the links, users won’t click.” The brief also describes an OpenAI researcher sharing a “hack to get around nytimes paywall” to Greg Brockman, who replied, “ah nice.” Per 404 Media, these statements come from materials the companies asked to keep sealed or redacted, underscoring the case’s core clash: the NYT alleges uncompensated extraction and substitution, while Microsoft and OpenAI maintain training on scraped web content is fair use.

The discussion splits between the practical degradation of information provenance and the structural economics of the AI transition.

On the technical front, a user’s anecdote about an LLM perfectly absorbing an original linguistic concept—only to hallucinate false citations when asked for the source—anchored a debate on attribution. While defenders pointed out that current models inherently lack document recall by design, critics argued that deploying such systems as search replacements acts as a deliberate "shell game." The crux of this camp's frustration is less about lost intellectual property and more about the "corruption of truth": models confidently stripping original work of its context and parroting distorted versions with an air of authority.

A separate thread zoomed out to the labor economics of automation. One prominent critique highlighted the irony of the tech industry—which spent a decade celebrating "software eating the world" and disrupting legacy sectors—now crying foul when cognitive labor becomes the target of standard corporate cost-cutting.

Commenters broadly dismissed Microsoft’s internal hand-wringing as hypocritical, though one user clarified that the quoted "doom loop" memo is actually from January 2023, immediately following ChatGPT's public launch. Suspicion toward the company remains high, with users speculating that Microsoft might eventually leverage its enterprise footprint to surreptitiously harvest corporate IP for ongoing model training.

NASA-IBM Lunar Foundation open-Source Geospatial AI Model

Submission URL | 53 points | by noobplus | 6 comments

Open-sourcing a lunar geospatial foundation model gives researchers and engineers a shared baseline for analyzing Moon data and building downstream tools, instead of training bespoke models from scratch. Backed jointly by NASA and IBM, the release lowers integration friction for geospatial workflows and makes auditing, extension, and reuse possible across academia, industry, and the open-source community.

The thread centers on the semantic distinction between genuine "open-source" AI and "open-weight" models. Commenters praise this release for including training methodologies and data catalogs, avoiding the "inscrutable binary blob" nature of AI models that only release their weights. This spawned a brief tangent on software control, with one user arguing that the real modern divide isn't open versus closed source, but local execution versus SaaS—asserting that even a compiled, closed-source local binary is preferable to an untouchable cloud service. Direct links to the model's Hugging Face repositories were also surfaced to bypass the corporate article.