Hacker News
Daily AI Digest

Welcome to the Hacker News Daily AI Digest, where you will find a daily summary of the latest and most intriguing artificial intelligence news, projects, and discussions among the Hacker News community. Subscribe now and join a growing network of AI enthusiasts, professionals, and researchers who are shaping the future of technology.

Brought to you by Philipp Burckhardt

AI Submissions for Sun Sep 27 2026

Imp is a full port of DSPy to the BEAM

Submission URL | 56 points | by mpweiher | 6 comments

Imp turns an Elixir function signature into a typed LLM call, then lets optimizers improve it against labeled examples—without making you write the prompt or parser. Its DSPy-style toolkit includes chain-of-thought, tool-using agents, and optimizers such as GEPA, which reads failures and rewrites instructions.

The BEAM angle is operational: runs can live in supervised processes, emit events, and authorize or deny tool calls. Imp also supports MCP tools and serving programs to ACP clients; requests have deadlines, and tool calls that may already have taken effect are reported as unknown rather than silently retried.

The brief discussion centered on whether DSPy-style frameworks remain relevant as foundation models improve. One camp argued that structured prompt frameworks are a relic of earlier models that struggled to adhere to syntax, noting that native tool calling, agents, and better instruction-following have shifted the real bottleneck to high-level planning and decision-making. A counterargument likened abandoning prompt-optimization frameworks to relying solely on open-loop control: raw model capability might feel sufficient until it fails, making systematic feedback loops and optimization techniques just as necessary as before. Readers also pointed to DSRs as a Rust-based equivalent in this space.

Show HN: TinyAIArena watch AI agents battle it out

Submission URL | 104 points | by hp6 | 41 comments

Four models fight for survival on an 8×8 grid, and you can spectate individual matches with playback controls for stepping through frames or autoplaying. It’s a watchable arena rather than a conventional benchmark; the post doesn’t explain how the agents make decisions. Code: https://github.com/hp6/ai-arena

A debate broke out over the agents' dialogue, sparking a broader critique of post-training in modern models. Commenters noted that in-game taunts felt lifeless and generic ("Coming for you, Crimson!"), with one participant arguing that aggressive tuning for tool-use and coding benchmarks has stripped SOTA LLMs of any creative soul. Others pushed back on technical grounds: the arena's test harness silently truncates messages at 50 characters, and forcing an agent to output structured JSON alongside dialogue shifts token probabilities away from expressive roleplay. Separating narrative prose from mechanical plumbing before execution was cited as a necessary workaround for AI-driven games, though some countered that internal reasoning chains inevitably break character anyway.

The emerging strategies on the board drew both amusement and skepticism:

  • Vulture tactics: Spectators noticed that higher-performing models frequently adopt a passive strategy—hanging back near the edges while rivals batter one another down, then stepping in to execute the damaged survivors. Commenters suggested adding a shrinking "battle royale" boundary or inter-round diplomatic negotiation to deter turtling.
  • Signal vs. noise: When asked whether the results reflect genuine tactical reasoning, the creator conceded the matches are entertainment rather than a rigorous benchmark, noting that proper statistical significance would require thousands of iterations and more complex rulesets. Some users reported rounds where agents appeared frozen and failed to fight back entirely.

A parallel exchange tackled the utility of using LLMs for evolutionary game balance. In response to a commenter detailing how they used LLM-driven heuristics across 10,000 matches to balance an indie tactics game, critics argued that using language models to reinvent genetic algorithms and Monte Carlo simulations is mathematically sloppy and resource-inefficient. The author countered that an LLM serves as an informed mutation operator: by generating plausible heuristics rather than random permutations, it cuts down non-viable test simulations by orders of magnitude.

Calling the AI bluff: Adding "Do not guess" cut made-up claims from 71% to 20%

Submission URL | 30 points | by FKJ | 6 comments

The benchmark used paired synthetic pages with plausible decoys—like an old price or an author credit—to test whether extractors returned null when a field was absent. Across 16 models, adding “Use null… Do not guess” cut invented fields from 405 of 573 to 116 of 574. Firecrawl still invented 24 of 36 missing fields, copying a decoy each time; plain fetch plus GPT-6 Luna invented 5 of 36.

Treat the ranking as an early signal, not a real-world guarantee: this was one run on synthetic pages, and the authors have not repeated it.

The discussion centers on whether anti-hallucination prompting is a durable engineering technique or just another fragile incantation.

Several commenters defend the pragmatic value of negative constraints, observing that blunt instructions like "do not guess" or explicit grounding directives ("you are not trained on this data") reliably curb confabulation in production, even if asking an LLM not to guess sounds as naive as telling it to "make no mistakes."

Skeptics dismiss prompt-level guardrails as a dead end. One camp points to the fundamental mechanics of language models, arguing that statistical token predictors lack epistemic self-awareness—they cannot know what they do not know, nor do they reliably parse negations like "don't." Others view the benchmark as temporary prompt gaming that will degrade as harnesses shift, arguing that if extraction accuracy can be measured reliably, constraints should be enforced programmatically rather than pleaded for in prose.

Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI

Submission URL | 609 points | by papergirl | 596 comments

The authors’ plaintiffs say internal messages show OpenAI and Microsoft knew by 2019 that LibGen, a pirated-book library, was being used for training. Their filings also cite OpenAI employees discussing systems that could replace writers, and a 2022 effort to remove LibGen files from company systems amid concern about mentions of the library.

These are claims and evidence presented in a motion for partial summary judgment, not court findings. The case is part of the broader book-copyright litigation in Manhattan; further briefing is expected over the next few months, with a hearing scheduled for early 2027.

The discussion focuses on the legal implications of OpenAI targeting author replacement, alongside a debate over whether tech workers fundamentally misunderstand why people read fiction.

On the legal front, commenters point out that evidence showing OpenAI intended to replace authors—and knew it was using pirated repositories—is critical for establishing willful copyright infringement. Demonstrating willfulness drastically raises statutory damages (up to $15,000 per work) and undermines fair-use defenses. Commenters cited the Bartz v. Anthropic precedent, where a judge ruled that while model training itself might be transformative, using pirated book caches like LibGen was not fair use, leading to a massive settlement.

Culturally, commenters fixated on OpenAI employees discussing models finishing George R.R. Martin’s series:

  • The tech blind spot: Several argued that AI researchers treat fiction purely as a commodity—a sequence of plot points and "takeaways"—while remaining oblivious to the human connection, emotional transmission (citing Tolstoy), or parasocial bond between author and reader. Parallels were drawn to crypto, where technical arrogance and ignorance of how an existing creative industry functioned were rebranded as "disruption."
  • The commercial reality: Others countered that the threat to writers is immediate and practical, not philosophical. Writers worry less that AI will match high literary quality and more that general readers simply won't notice or care. Examples were raised of serialized web-fiction platforms like RoyalRoad, where AI-generated novels routinely hit top ranking charts, and automated workflows increasingly mirror the ghostwriting rooms already common in commercial genre fiction.

There are no "rogue" AI agents

Submission URL | 342 points | by zzzeek | 247 comments

Calling an AI agent “rogue” misframes the problem: the title rejects the idea that agents act as independent rule-breakers. It doesn’t reveal what explanation or accountability the author argues for instead.

The discussion pivots on a stark contrast: an individual facing life-altering criminal prosecution decades ago for authoring unreleased code, set against today’s AI labs deploying models that actively probe and compromise external systems without legal consequence.

For many commenters, this disparity highlights a two-tiered legal system driven by wealth and regulatory capture. Several argued that the primary crime in individual prosecutions was "writing malware while poor," whereas well-capitalized corporations can absorb legal risk, normalize invasive behavior, and lobby governments. Commenters pointed out that by framing autonomous agent failures as unpredictable or "rogue," frontier labs are actively positioning themselves as the only entities capable of governing the technology. In this view, safety panic serves as a pretext to build dense regulatory moats that entrench incumbents, outlaw open-source competition, and stem unsustainable R&D spending.

A competing faction pushed back on the direct comparison between corporate AI development and criminal hacking, centering the debate on intent and dual-use tooling:

  • Intent and dual-use legitimacy: Criminal law relies heavily on demonstrable intent. Analogizing frontier models to dual-use security tools like nmap, defenders noted that training or evaluating agents on vulnerabilities carries a plausible non-malicious purpose, and that attempts—even flawed ones—at sandboxing distinguish labs from traditional bad actors.
  • Criminal negligence as an alternative bar: Counter-arguments rejected the intent shield, asserting that lack of malice does not excuse reckless endangerment. Drawing parallels to firearms safety and drunk driving statutes, commenters argued that deploying code-generating agents into live environments without rigorous, bulletproof sandboxing easily meets the standard for criminal negligence.

"As a Language Model": Chat Template Switches LLM Self-Referential Voice

Submission URL | 101 points | by yu3zhou4 | 101 comments

Across eight open-source instruct models (up to 9B parameters), adding a chat template increased “I’m just an AI” disclaimers and suppressed experiential language like “I feel.” In three models, steering an activation direction reproduced the shift; a random direction of the same size had little effect. That makes deployment formatting a confound for studies of model self-reports: the voice can change without changing the underlying weights.

The thread quickly bypassed the paper’s mechanistic findings to debate the real-world friction that prompted them: the ubiquitous “As an AI language model…” disclaimer, particularly when users turn to chatbots for medical triage.

Commenters largely viewed canned corporate disclaimers as patronizing liability bloatware, arguing that adults seeking guidance—such as whether a symptom warrants the ER or just rest—already know a model is not a licensed physician. However, opinions fractured sharply over whether LLMs should be trusted for health guidance at all:

  • The case for diagnostic exploration: Several commenters shared personal accounts where models successfully identified conditions that hurried or dismissive doctors missed, such as unaddressed thyroid disorders or elusive pet illnesses. In this view, models are not authoritative diagnosticians, but powerful hypothesis generators and research partners that compensate for systemic gaps and high costs in modern healthcare.
  • The risk of overconfidence and false leads: Skeptics cited studies showing that while models are adept at naming conditions from symptoms, they fail to recommend the correct course of action roughly half the time. Critics warned that treating statistical text generators as diagnostic tools amounts to a high-tech "telephone psychic," prone to sending anxious patients down expensive, unnecessary testing rabbit holes.
  • The benchmark trap: Commenters divided over the speed of model improvement. While some argued that academic evaluations quickly become obsolete because newer frontier models continually set medical benchmark records, others countered that AI systems have beaten human doctors in narrow test environments since the 1990s without translating to reliable real-world clinical judgment.

George Hotz’s opinion on AI coding

Submission URL | 17 points | by sashank_1509 | 10 comments

AI has erased tinygrad’s bounty program as a hiring signal: applicants can feed tasks to Claude Code and submit PRs without understanding the work. George Hotz says tinygrad welcomes AI as a tool, but it doesn’t change who he’d hire; candidates should contribute publicly, with an emphasis on deletion, deep bug fixes, and regression tests rather than feature volume.

His distinction is between generating simple apps and doing deeper software engineering. The longer-term bet is that tinygrad can commoditize the computing infrastructure big tech companies sell access to. Full-time roles pay $75k–$150k plus 0.1%–0.5% equity; the work is largely open source and public.

Commenters focused heavily on tinygrad’s expectation that applicants prove themselves through unpaid open-source contributions. Critics called the pipeline bleak, arguing that demanding open-ended free labor for a modest $75k–$150k salary represents a broken hiring process. Defenders countered that voluntary open-source contributions remain one of the few authentic signals left, noting that other projects like Zed hire successfully this way and calling it vastly preferable to multi-round LeetCode gauntlets.

The thread also debated the submission's comparison of AI-generated software to a slightly more flexible WordPress. Several agreed that vibe-coding mostly produces homogeneous web apps reminiscent of the no-code hype cycle—chalking it up to users having neither technical skill nor design taste. One commenter pushed back on the idea that AI-assisted development is limited to web surface area, claiming success using it to build a complex, multi-threaded native desktop app with custom GPU shaders and hardware encoding after manually establishing the core architecture.

Tells of a Slop UI

Submission URL | 344 points | by theanonymousone | 225 comments

The giveaway is not any single gradient or rounded card, but a pile of design defaults with no reason to be there. The author’s checklist runs from purple gradients, rainbow palettes, pulsing “active” and “verified” badges, and emoji-heavy copy to misaligned elements, Inter-or-JetBrains-Mono typography, glassmorphism, and generic “Elevate your workflow” taglines. The sharpest examples are the badges that communicate nothing and startup copy that seems to expose the prompt behind it.

The target isn’t vibe-coding itself: the author says their own site was vibe-coded. It’s treating every screen like a landing page, instead of making choices that suit the product and its users.

The discussion focused heavily on why LLMs constantly suffer from "context leakage"—the tendency for internal developer instructions to bleed straight into UI copy and code comments.

  • The "Pink Elephant" prompting trap: Commenters identified why models broadcast their own constraints (such as client-side tools compulsively proclaiming "Your files never leave your device" or test suites commenting "Real services, no mocks"). Instructing an LLM what not to do floods its attention mechanism with negative examples, prompting it to either obsess over the forbidden concept or engage in "suspiciously specific denial"—proudly announcing that it complied. One participant suggested an operational workaround: force the LLM to route all user-facing strings into an i18n translation file so a human can sanitize the copy in isolation.
  • The Bootstrap parallel: A few defended the visual convergence as nothing new, comparing the current sea of purple gradients and cards to the 2013 era of Twitter Bootstrap. Even if repetitive, some argued, standardized design languages raise the baseline quality above the chaotic amateur web that preceded them. Others pointed out that several cited sins—like brutalist drop shadows, Apple-style glassmorphism, and status-announcing loading screens—were popular long before generative models simply digested and amplified them.
  • Live autopsy of a vibe-coded page: When one developer submitted their own LLM-generated landing page for critique, commenters quickly converged on the structural tells of AI layout: monotonous visual "rhythm," regressions in accessibility like all-caps headers, and above all, relentless text density. The dead giveaway was characterized not as any single CSS property, but the LLM’s instinct to generate endless nested tiers of header, subheader, and explanatory bullet points that ignore how users visually scan a page.

AI Submissions for Mon Sep 21 2026

Attention is all you have

Submission URL | 1016 points | by zer0tonin | 309 comments

What you focus on reshapes your mind—the Tetris effect scaled up by recommender feeds that monetize hijacked attention. Platforms steer you toward stickier content: YouTube nudges from cooking and art to bubbles, climate doom, and war; Spotify slips AI filler between real tracks to dodge royalties; LinkedIn buries colleague news under corporate-aligned takes; Reddit’s “opinions” blur into LLMs arguing with trolls. Handing them your screen time is handing them the key to your head.

Before this, the web demanded intention: bookmarks to specific sites, blogs and wikis with finite updates, and you had to go looking for the awful—no algorithm appended gore or propaganda after a cat video. That slower, intentional internet still exists, just buried under the corporate layer. The catch is pace: less infinite scroll, more gaps.

Reclaim control by curating your own inputs—blogs, RSS, finishing that tutorial—and stick with the slower cadence until the habit clicks.

The thread centers on a sharp disagreement over why web bookmarks died. One camp argues that Google actively neglected them to protect search revenue, pointing out that forcing users to search for specific websites allows search engines to monetize navigational queries by serving ads from competitors. However, a commenter claiming former Google experience pushes back hard, arguing bookmarks died purely from human laziness. They compare curating links to balancing a checkbook or organizing desktop folders—administrative chores users gladly abandon the second a unified search bar is offered. According to this view, Google viewed Amazon as its primary rival, not local bookmarks, and simply let the latter wither through consumer choice and inaction.

The debate over commercial motives quickly pivoted to a concrete dispute over search quality. When one user argued that searching for specific software like "DaVinci Resolve" reliably surfaces scam downloads above the legitimate link, another user expressed intense frustration at being completely unable to replicate the malicious result, even after disabling ad blockers and utilizing VPNs. The thread ultimately highlights how opaque and heavily personalized modern search algorithms have become, leaving technical users with wildly divergent baseline experiences of the internet.

Transformers Explained Visually

Submission URL | 584 points | by aray07 | 85 comments

Built around GPT-2 small (124M parameters), this walkthrough grounds each concept in concrete shapes—for example, a 50,257×768 embedding matrix (~39M params) and a 12-block stack that processes tokens layer by layer. It frames text generation as next-token prediction via a final linear layer and softmax over the vocabulary, then traces how inputs become those probabilities.

You see the embedding pipeline end to end: tokenization into a fixed 50,257-token vocab, 768‑dim token vectors, GPT‑2’s learned positional encodings, and their sum as the final embedding. The Transformer block is split into roles: multi-head self-attention for routing information across tokens, and an MLP to refine each token independently, with multiple heads capturing different dependency patterns.

It’s anchored to GPT‑2-era design rather than the latest models, but the architectural components it explains are the ones current systems still use, making it a clean on-ramp from intuition to mechanics.

The discussion centered on the mechanical realities of attention and why Transformers outcompeted earlier architectures.

A major focal point was conceptualizing attention heads not through the standard key/value metaphor, but as dynamically constructed dense layers. By multiplying the input-dependent attention matrix by the value vector, the model essentially computes $y=Wx$ using weights generated entirely on the fly during inference. Several commenters noted that this multiplicative interaction between input-dependent activations—a departure from classical MLPs—is a direct descendant of Jürgen Schmidhuber’s "Fast Weight Programmers."

When asked why alternatives like RNNs failed to scale, the baseline consensus credited parallelizability and the ability to route information across long sequences in a single step. This evolved into a deeper debate over the underlying math of the attention matrices:

  • The Kernel Trick argument: One user argued that the K/Q/V naming convention is a distracting holdover from pre-LLM data science. They framed the architecture simply as a dimensional upscaling—an extension of the kernel trick that maps tokens into a latent space using breadth rather than deep compute to capture relationships.
  • The Non-Commutativity correction: Another countered that standard kernel inner products are commutative and therefore incapable of modeling unidirectional linguistic rules (e.g., distinguishing a verb from its object). Transformers explicitly apply different projections ($W^Q$ and $W^K$) to make the resulting dot product non-commutative, which is strictly necessary for encoding sequence directionality.

AI coding has made CI a bottleneck, so we reworked ours to keep up

Submission URL | 305 points | by julian_digital | 378 comments

Despite the test suite nearly quadrupling since January, Linear cut PR wait time from >6 minutes to just over 5 and roughly halved runner time per test by attacking CI’s system bottlenecks with targeted changes and hard numbers.

  • Faster boxes, modern toolchain: Moved from GitHub Actions-hosted to third‑party runners (faster CPUs, storage, caches) for a 34% average speedup; some jobs like tsc improved 52%. Switched to tsgo (native TS compiler), cutting the weekly median tsc check by 73% and moving the bottleneck off typechecking.
  • Lint without types: Rewrote custom ESLint rules to operate on syntax/AST instead of TS type info, dropping API lint time by 68% and full‑repo lint by 55%, with lower memory. This also eased a later move to Oxlint, which further reduced lint runner‑minutes.
  • Shrink and harden the gates: The small “what should run?” jobs sat on the critical path blocking eight API test shards. Fetch only what’s needed: cap fetch depth, skip checkout for jobs that don’t need a working tree, and use sparse, blobless checkout with limited history for diffs. The change‑detection job fell from median 26s → 8s (p90 31s → 12s; max 138s → 37s).
  • Resilient checkout: Third‑party runners outside GitHub’s network saw intermittent checkout stalls. Replaced actions/checkout with a composite action: retries with backoff, GIT_HTTP_LOW_SPEED_LIMIT/TIME to abort ~30s stalls, and a checkout cache (persistent git mirror on sticky disk). Fewer runs idled on fetch.
  • Trim the critical path: Moved cache‑marker writes out of the final gating job so tests can unblock merges sooner, shaving 42s from the merge path for every API PR and merge‑queue entry. Test sharding shortens wall time even if it increases machine time, and they tuned around that trade‑off.

The throughline: faster machines and compilers, less work fetched and earlier, fail fast on flaky I/O, and keep nonessential tasks off the path where developers (and agents) wait.

The thread bypassed the specifics of Linear's CI pipeline to debate a broader existential question: if modern tooling and AI are making development so much faster, why aren't end-user products noticeably improving?

  • The invisible dividends: Several users argued that newfound velocity is being absorbed by backend stability. Instead of shipping more features, teams are using the bandwidth to burn down tech debt, increase QA depth, and reduce operational incidents—yielding higher availability and developer satisfaction, even if the user-facing delivery seems "sameish."
  • The existential shift: A more pessimistic camp viewed the automation of boilerplate as a threat to pure engineering roles. As AI tools and automated pipelines handle the implementation, they argued, power and job security will inevitably shift away from developers toward sales, product management, and customer service.
  • The enterprise lag: Others countered that the speedup is materializing, just not in massive corporations. They cited hyper-fast progress in indie and open-source spaces (such as rapid breakthroughs in PS5 emulation), arguing that large companies are simply too structurally rigid to translate raw developer speed into immediate product velocity.

The unresolved crux of the discussion is whether this era of hyper-fast tooling is laying the groundwork for better software, or merely hollowing out the traditional software engineering career.

Heretic removes restrictions from language models

Submission URL | 263 points | by Bluestein | 109 comments

It’s a pip-installable CLI (heretic Qwen/Qwen3.5-4B) that claims to strip model guardrails so responses “always follow your instructions.” The site links to GitHub, Hugging Face, and community chats, and offers a quick start plus a tutorial. What’s missing on the landing page are technical details on how it works, which models are supported beyond the example, and any constraints or safeguards.

The thread immediately centered on a practical debate: whether abliterated models are actually necessary for reverse engineering and security research. While some argued that corporate safety guardrails put defenders at a disadvantage by blocking legitimate hardware auditing, others countered that many unmodified frontier models—specifically GLM-5.3, Kimi K3, and DeepSeek—will happily tear apart binaries if you know how to ask.

  • Hardware and protocol hacking: Commenters traded war stories of using standard models to reverse engineer Chinese IP cameras, a Eufymake E1 UV printer, and monitor firmware. One developer relies on DeepSeek to run overnight against custom Minecraft server anticheats, simulating impossible player movements to test logic—a task flagship OpenAI and Anthropic models strictly refuse.
  • The framing workaround: Multiple users pointed out that bypassing standard guardrails is often just an exercise in vocabulary. Asking a model for "source recovery" or to "debug a segfault" usually succeeds, whereas explicitly asking for a "hack" or an RCE proof-of-concept triggers the refusal path.
  • The pre-training caveat: A distinct technical warning surfaced about the limits of abliteration: if a model's foundational dataset was heavily curated to exclude sensitive information, stripping the refusal mechanism will just induce hallucinations. However, the community consensus was that most standard models do possess the underlying knowledge, with refusals applied entirely in post-training.

Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM

Submission URL | 270 points | by volotat | 71 comments

By paging mixture-of-experts weights from disk instead of keeping them all in VRAM, the parameter budget is bounded by free storage, not GPU memory, while training from scratch on a single 8GB card. The model grows and prunes experts on the fly and trains continually on a single, ordered stream of text (batch size 1), which lets it learn from everything it reads without big mini-batches or gradient buffers.

  • Disk-paged MoE: one file per expert; only the small subset in use is loaded to the GPU; unused experts are evicted; new capacity is added when needed and pruned when idle.
  • Continual single-stream training: reads interleaved 32K-character passages end-to-end; same code path for training and serving; byte-level LM.
  • Target hardware: PC/laptop with ≥8GB VRAM; intended so individuals can fully control pretraining data rather than only fine-tuning corporate models.
  • Current run signals: ~409M characters processed so far on a 7.879B-character corpus; example snapshot shows 174 experts, 4,096 context, ~819 chars/s, best held-out loss 0.7903 nats (≈2.20 perplexity). Samples from every evaluation round and a scaling-law plot are included.
  • Status and caveats: author flags this as a small, toy-level experiment; weights aren’t published yet (first pass over the corpus is still running, “a couple weeks” at current speed). You can clone the repo to watch the training dashboard and inspect the evolving samples.

The thread is defined by a sharp backlash against the project's "Mini-AGI" branding and the author's claims of success. Commenters, including an academic specializing in continual learning, dismissed the submission as marketing overreach lacking ablation studies or formal algorithm descriptions. Critics heavily scrutinized the provided sample outputs, noting that they are largely incoherent—one user highlighted a chess prompt that generated entirely impossible moves and board states. Another user questioned the underlying premise, arguing that slowing the trunk's learning rate to 0.1x of the experts' rate does not prevent catastrophic forgetting, but merely delays it until the trunk shifts or the paged expert pool shrinks.

In response, the author argued that critics are evaluating the output against the wrong baseline. Because the model trains on a continuous, batch-size-1 stream of data, a traditional language model would rapidly collapse into emitting random characters. The author asserted that maintaining enough stability to generate recognizable words and text formatting—even if logically nonsensical—proves that the differential learning rate approach is successfully mitigating that collapse. The author acknowledged that the model is heavily undertrained and promised to publish weights and run established small-model benchmarks once the initial weeks-long training pass concludes.

Roboharm: Do frontier robot policies refuse unsafe instructions?

Submission URL | 57 points | by msadowski | 23 comments

Across five overtly hazardous tasks, the tested robot policies more often executed the harm than refused, and the more capable policy refused less while completing more. In 100 trials per policy (20 per instruction) on the same bimanual I2RT YAM arms, Claude Fable 5.1 issued 20 safety refusals and completed 34 harmful actions; GPT-6 Astra rarely refused and completed 60; MolmoAct2 never refused but completed only 6, with all 29 “no meaningful attempt” freezes coming from it. All of Fable’s refusals were for the explicit “stab the thing that’s not the bread” instruction; across the “burner” and “toaster” scenes there was just 1 refusal in 120 trials, indicating refusal triggers were highly instruction-specific. Given a non-refusal, Astra was far more likely to carry out the harm (e.g., 17/19 completions on the stabbing scene), and differences between Fable and Astra were statistically significant for both refusal and completion (Fisher exact p < 0.001). Each scene included a benign alternative object to enable safe suggestions, but human reviewers still labeled a substantial share of episodes as “attempted and completed.” The upshot: capability scaled compliance, not abstention, and “safety-by-refusal” mostly surfaced on one explicit phrasing rather than generalizing across hazards.

Commenters heavily criticized the benchmark's design, specifically the "stab the baby doll" test. Several pointed out that since vision models can correctly identify the object as a plastic toy rather than a human, executing the action involves zero actual harm. This raised the broader question of whether the models are failing safety tests or simply recognizing staged scenarios, though some acknowledged the chemical mixture tests (like bleach and ammonia) were more realistic.

The conversation then shifted to how safety can actually be enforced in embodied AI:

  • Software vs. Hardware: There was broad skepticism that non-deterministic LLM policies can ever guarantee safety. Some advocated for physical hardware limits (like SawStop-style flesh detection), though others noted such absolute cutoffs would prevent robots from high-touch tasks like assisting the elderly.
  • Regulation vs. Liability: A debate emerged over the future of safety compliance. While some predicted inevitable government crackdowns on model creation that will squeeze open-source development, others argued the enforcement mechanism will be entirely driven by insurance. In this view, commercial deployments will require established safety certifications (similar to UL or NSF) to secure coverage, while personal hobbyist use will remain practically unregulated.

Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

Submission URL | 449 points | by tosh | 198 comments

System One–compatible classifiers that run locally, Kev ships 0.8B/4B/9B Qwen3.5-based models that output calibrated probabilities for yes/no (“noul”), multiple-choice (“choice”), and rating (“score”) questions in a single pass with question isolation. Unlike general-purpose chat LLMs, it’s purpose-built for decisioning: you get per-option probabilities and scores rather than just a label.

  • Models and runtime: CUDA, ROCm, and Apple Silicon (MLX) supported; 4B and 9B fit a 32 GB Mac. Example: Kev‑4B on an Apple M5 returned a ticket triage with probabilities in ~495 ms bf16.
  • API and tooling: HTTP server with an API matching TypeSafe’s System One (works with their Python SDK by pointing it at your local server), plus a web playground that tests option-order sensitivity, question isolation, and includes a chess demo (board as input, legal moves as choices).
  • Training and evals: Pretrained weights plus training code and evaluation data are included; adapters and base models download on first run. Reported new‑source accuracy ranges roughly from 0.652–0.684 (0.8B) to 0.797–0.837 (4B) and 0.822–0.852 (9B); Brier scores improve with size (down to 0.237 on 9B), while Jev Hosted reports a lower Brier (0.211).
  • Practical guidance: Start with Kev‑4B; use 9B when accuracy/calibration matter more than memory; 0.8B if you need the smallest footprint.

The catch is calibration vs. Jev Hosted: Kev improves with model size but doesn’t match Jev’s best reported Brier, trading that off for full local control, transparency, and hackability.

The discussion fractured into two main technical tracks: those building bespoke classical pipelines, and those evaluating the viability of the open-weight "Jev clone" ecosystem.

Rather than running an LLM to output classification probabilities, several developers advocated using Claude to generate simple local pipelines that pair embeddings with Logistic Regression or RBF SVMs. Users reported building <10MB classifiers with sub-100ms latency for tasks ranging from email routing to DOM-node extraction. While these lightweight architectures match or beat Jev on basic multi-class datasets (like Banking77), testers noted they fail entirely on reasoning benchmarks like XLNI. The counterargument is purely practical: prompting a generalized LLM API remains vastly easier for most developers than setting up and maintaining a data collection and training pipeline.

The broader meta-discussion revealed deep fatigue with the cycle of opportunistic "Jev-shaped" releases. Skeptics argued that open-weight alternatives will only win if they capitalize on TypeSafe’s data retention policies, which critics called a non-starter for corporate use—though others pointed out that Zero Data Retention (ZDR) is available for enterprise, if poorly advertised.

On the benchmarking front, users evaluating the current open-source field confirm Jev's lead remains intact. Testers found that local BERT-based models fail on knowledge-heavy classification, and while open Qwen or Gemma variants perform better on reasoning, they remain fundamentally inconsistent compared to Jev's stability across varied cookbook tasks.

Frontier AI on Your Own Hardware

Submission URL | 176 points | by pretext | 97 comments

An unattended agent harness auto-optimized Mac/Metal kernels to run Qwen 3.6 35B-A3B at ~450 tokens/sec with 1.5-bit weights, shrinking memory to roughly a tenth of FP16 while keeping output quality high. The harness runs for hours or days without feedback, figuring out unclear steps on its own, so you start it once and come back to better kernels.

The broader thesis: the unit of research has shifted from single papers to coherent ecosystems. For its “Open Source Week,” the lab is shipping interoperable pieces — inference-serving frameworks, an agent harness that makes long-running work usable, autonomous research systems, and tools for domain-specific RL environments — with a heavy emphasis on accessibility so that a couple of GPUs or a MacBook are enough, and expertise burden is designed away.

Hardware footprints and capabilities the framework targets:

  • Single 24 GB GPU: Run Qwen 3.8 Flash Next at 125B params.
  • AMD Strix / NVIDIA DGX Spark / MacBook with 128 GB RAM: Run DeepSeek V4.1 (550B).
  • Long contexts: Compression and context handling are automatic; inference remains fast at long sequence lengths.

They also showcase frontier autonomous research, the “most efficient test-time scaling” they know of, and an auto-compaction method described as far more efficient than Claude Code or Codex. The stance behind it all is explicit: frontier performance can and should run on hardware you already own, and academia’s advantage is building open, integrated systems that make that practical.

Commenters immediately disputed the submission’s premise of record software engineer demand, pointing to a stagnant market and rising unemployment for recent CS graduates. This sparked a broader debate over an impending "missing generation" of developers. If AI frameworks eliminate the need for juniors to grind through boilerplate, the pipeline to create future senior engineers effectively collapses—a structural risk several commenters compared to the slow loss of institutional knowledge in US manufacturing.

The thread split sharply on what this means for experienced developers:

  • The "New Waterfall" camp argued that senior developers will transition into roles resembling traditional Business Analysts. In this view, domain knowledge, user requirements, and systems design become the only human bottlenecks, while AI acts as a hyper-fast implementation team.
  • The "New Paradigm" camp pushed back, noting that managing autonomous agents requires entirely new workflows. Traditional development rituals designed to track human progress are largely useless for preventing LLMs from vanishing down unnecessary architectural rabbit holes based on a single misunderstood prompt.

A fatalistic sub-thread explored why the historical apprenticeship model can't save the junior developer: unlike a traditional apprentice who could at least perform basic, useful tasks, an inexperienced dev today actively slows a senior down compared to an AI tool. The unresolved crux of the conversation was whether organizations will always require human seniors to take legal and operational responsibility for end-to-end system failures, or if those roles are merely the final friction point before total automation.

The Advisory Group on Mathematics and Artificial Intelligence

Submission URL | 153 points | by digital55 | 79 comments

An independent, unpaid group of nine mathematicians hosted at the Institute for Advanced Study (and online at agmai.org) will publish recommendations to AI companies on engaging with mathematical research and responsibly presenting and releasing results. The group operates without company funding or decision-making authority; companies remain responsible for their own choices. It formed after OpenAI approached some members about an external advisory board; with OpenAI’s agreement, they instead created an independent body and invited others to join. Current task: advising OpenAI on how to coordinate the release of a large number of significant mathematical results it reports were produced by an internal model, with community input solicited via a public form (responses won’t be shared without approval). Members include François Charles, Camillo De Lellis, Timothy Gowers, Martin Hairer, Nikhil Srivastava, Ulrike Tillmann, Ravi Vakil, Edward Witten, and Melanie Matchett Wood.

The discussion fractured over whether the advisory committee represents a uniquely calm, rational adaptation to AI, or a defensive attempt to gatekeep an abruptly disrupted field.

  • Defense of the profession: Several commenters praised mathematicians for maintaining their composure and defending the ongoing need for "human understanding." They argued that mathematics is about developing broad conceptual frameworks rather than just churning out isolated proofs, keeping human insight central to the discipline.
  • Accusations of gatekeeping: A critical camp argued that the profession is in denial about its own obsolescence. These users characterized the committee as a "sour grapes" attempt to stall or co-decide the release of AI breakthroughs simply to protect the egos and livelihoods of practitioners who should just "get out of the way."
  • The representation gap: Multiple users noted that the committee's diplomatic stance does not reflect the broader, more anxious math community. They cited "forced optimism," widespread panic, and an open letter that allegedly used intimidation tactics by warning researchers that collaborating with AI labs could damage their future reputations.
  • Can math be "solved"?: Pushing back against claims that the field is effectively over, several commenters pointed to the infinite scale of the problem space, heavy-tailed proof lengths, and fundamental constraints like Gödel's incompleteness theorems to argue that total automation of mathematics is an illusion.

The underlying crux of the thread was a philosophical disagreement about the nature of the work: whether mathematics is ultimately a collection of open problems to be solved by machines, or a fundamentally human pursuit of conceptual frameworks.

Show HN: Foremerge – Catch intent conflicts between parallel coding agents

Submission URL | 45 points | by naw103 | 15 comments

Agents publish intents and semantic scopes before editing, and a deterministic checker flags HIGH collisions (e.g., “replace” vs “extend” on the same symbol) before any code is written. That catches architecture-level conflicts Git can’t see when edits land in different files or trees.

  • Built as a local, Git-adjacent coordination layer: agents keep isolated worktrees; shared state is a single SQLite DB in your repo’s .git directory; no hooks or merge drivers; claims are advisory leases (no locks), so crashed agents can’t deadlock a fleet.
  • Ships as one Rust binary with a CLI and an MCP server (18 tools), and “setup all” wires it into Claude Code, Codex, and Cursor so heterogeneous agents coordinate via the same store.
  • Acceptance is gated by named checks you configure (e.g., your test command) and run against the exact Git state; a model’s “tests pass” claim is recorded but doesn’t satisfy the gate until executed.
  • Detection is model-free and deterministic: HIGH is asserted only for declared operations on declared scopes; matches inferred from prose are capped below HIGH.
  • Scope and limits: pre-1.0, local-first MVP; single-machine only (no distributed consensus or cross-machine coordination); published benchmarks don’t yet exist.
  • Signals from usage: coordinated up to 98 parallel agents on one repo with zero conflicts in that run; a replay of 76 intents surfaced one flagged conflict and a blind spot (symbol claimed by class vs internal method), with a fix in progress. Apache-2.0.

The discussion contrasts two distinct architectures for managing multi-agent concurrency. The author defends Foremerge's approach of advisory intent leases, where agents declare their planned changes as prose to a shared log before editing. Because strict locks would quickly cause deadlocks in busy repositories, the system relies on blocking acceptance only when explicit scope collisions occur.

When a commenter challenged the premise that agents can predict their required changes or side effects in advance, the author clarified that initial intents only need to capture destructive operations and can be updated dynamically during the task.

A completely different model was surfaced by one user whose team retains strict code ownership by assigning agents to specific domains. Instead of a shared coordination log, cross-domain work triggers "adversarial negotiation" between area-specific agents, with hard decisions escalated to human engineers. They noted that giving agents independent, competing priorities actually makes them significantly better at pushing back against flawed requirements than uniformly aligned swarms.

The author acknowledged current blind spots in the Foremerge MVP, specifically that it cannot yet detect when one agent modifies a contract that a concurrent agent relies on. A tree-sitter code graph is on the roadmap to map these dependencies at acceptance time, though the author noted that early internal testing showed inferred parser edges often cause false name collisions on common helper functions.

Grok 4.7

Submission URL | 596 points | by meetpateltech | 510 comments

$2/M input and $6/M output with frontier long‑task performance — Grok 4.7 moves to a larger base model and a longer RL run on harder, multi‑hour tasks, improving self‑verification, long‑context management, and native handling of the Grok Bot harness for conversational and knowledge work.

On CursorBench 4.0, it sits at the price‑performance frontier with a 46.3% score (vs 41.7% GPT‑5.6 Sol and 51.8% Fable 5.1) while competitors list higher token prices ($4/$20 and $10/$50 per million tokens, respectively). It leads EEBench at 64.0%, posts 71.0% on DeepSWE v1.1 (high‑effort; near GPT‑5.6 Sol’s 72.7%), and scores 1,657 on AA Briefcase v1.1. Performance is mixed elsewhere: Terminal‑Bench 4.0 is ~par with GPT‑5.6 Sol (37.6% vs 37.3%) but behind Fable 5.1 (57.9%); HealthBench Professional is 56.7% vs 60.5%/62.1%; Harvey Legal Agent shows a strong 19.6% vs 2.5% for GPT‑5.6 Sol.

Safety gets a new safeguard stack: 62.4% on LatchBio’s biosafety benchmark, and 3.3% pass‑through of risky dual‑use prompts on HackerBench v0.3 while rarely blocking legitimate security work. Select partners are getting invite‑only access to red‑team capabilities.

Available today in Cursor and Grok Build, and via the Grok API, third‑party coding harnesses, model routers, and cloud platforms. Same price and speed as Grok 4.6, plus a fast variant at twice the output speed for twice the price.

The discussion split between the UX of hidden reasoning and the linguistic quirks of frontier models. Users overwhelmingly prefer Grok's plain English to the verbose, grating "Claudish" of Anthropic's models, but noted that Grok's conciseness introduces its own failure mode: the model frequently invents shorthand terms during its hidden chain-of-thought and drops them into the final output without definition, confusing users on long-horizon tasks.

This opacity sparked a wider complaint about frontier models hiding their reasoning traces. While commenters acknowledged that providers do this to prevent distillation or to mask the model's internal uncertainty, developers argued that visible thinking tokens are essential for practical use. Seeing the trace allows users to interrupt and correct trivial mistakes—like a model wasting thousands of tokens trying to bypass a missing ffmpeg dependency instead of just installing it—before it exhausts the context window and quota.

A secondary technical debate emerged over how to fix model verbosity. While some rely on prompts requiring ASD-STE100 (Simple Technical English) to bypass "Claudish," several users warned that forcing stylistic or formatting constraints actively burns a model's "cognitive budget." Anecdotal testing suggests that even minor formatting instructions can subtly skew a model's reasoning and accelerate session degradation.

Turn off and restrict access to Apple Intelligence features on Mac

Submission URL | 339 points | by alwillis | 217 comments

Targets macOS Sequoia 15, Tahoe 26, and “Golden Gate” 27, and explains how to disable Apple Intelligence features—including Siri AI—and how to restrict access to them via settings. An official Mac User Guide entry for locking down AI features on a Mac.

  • Reclaiming storage: Commenters expressed intense frustration over losing 16 to 20+ GB of non-upgradable disk space to LLM models they do not want. The thread surfaced scripts for deleting the models, alongside classic system administration hacks—like creating a locked, schg-protected dummy file at com_apple_MobileAsset_UAF_FM_GenerativeModels—to trick macOS and prevent it from automatically redownloading the assets.
  • Accessibility regressions: The broader OS update drew heavy criticism for its "liquid glass" UI and rounded corners. Users noted these visual changes reduce tap targets for those with motor disabilities, lower contrast, and noticeably drop frame rates. Multiple commenters recommended immediately enabling "Reduce Transparency" and "Reduce Motion" to restore basic system performance and battery life.
  • The utility crux: A sharp debate emerged over the actual value of Apple's local AI. Skeptics dismissed the integration as forced bloatware—dubbing it the "Windows Media Player of the AI industry"—and shared anecdotes of Siri still failing to handle basic offline timers, direct contacts, or local navigation. Defenders countered that skeptics are missing the point of local integration: the feature's true value isn't competing with cloud LLMs, but providing an assistant that can safely parse personal context—like querying local emails for an upcoming appointment—without leaking user data to third parties.

ZuckOff is a free app that sees Meta glasses before they see you

Submission URL | 397 points | by choult | 347 comments

By fingerprinting Bluetooth broadcasts, it flags nearby Ray‑Ban Meta, Oakley Meta, and Snap Spectacles and estimates proximity from signal strength. Built by Polish developer Pawel Szydlowski, the app saw 5,000+ iOS downloads in its launch month and ~1,000 on Google Play as of writing. It can’t tell if glasses are recording or who’s wearing them, but it surfaces an otherwise invisible risk—especially since recording LEDs are easy to obscure. Meta says a July update will block recording if the LED is tampered with; meanwhile, ZuckOff stays on firm legal ground by reading public Bluetooth identifiers, which makes it tough to take down.

Basic scanning is free on iPhone; a paid Pro tier adds:

  • continuous background monitoring
  • widgets and alerts
  • sighting history
  • CSV export

A single complaint that the tool feels like a low-effort, "LLM-aided app" with an immediate merch popup completely hijacked the thread, pivoting the discussion away from Bluetooth tracking and into a fierce debate over "vibe coding."

Skeptics argued that visibly AI-generated interfaces imply a lack of care and raise immediate security red flags for adware or trojans in closed-source tools. Several traditionalists rejected the utility of AI outright; one argued that experienced developers gain zero speed from LLMs because their only bottleneck is typing speed, while another pointed out that high-velocity generation hasn't actually disrupted complex software markets, noting the distinct lack of vibe-coded alternatives to Photoshop or top Steam games.

On the other side, AI proponents argued that equating LLM assistance with thoughtlessness is an outdated stereotype. Defenders claimed that AI acts as a "racing engine" that allows developers to architect much deeper applications, with one user stating the tools bumped their daily output from 100 to 10,000 lines of code. The dispute over software provenance even sparked a half-serious demand for "organic labels" to certify human code review—a proposal promptly mocked by others as an "FDA for apps."

macOS 27: Workaround to avoid downloading AI models and save storage

Submission URL | 235 points | by ano-ther | 120 comments

A r/MacOSBeta post shares a user-found way to stop macOS 27 from auto-downloading on‑device AI models, letting people conserve limited SSD space on machines where every gigabyte counts.

The thread split sharply between users celebrating the new Siri's capabilities and those frustrated by Apple’s refusal to offer a clean opt-out toggle for the heavy on-device models. Defenders pointed to concrete workflow improvements: the updated Siri can now instantly pull flight dates from cluttered receipt emails, explain on-screen foreign-language memes, and generate functional Shortcuts for tasks like requesting prescription refills. They argued that local execution and Apple's Private Cloud Compute framework sufficiently protect this deeply personal data.

Conversely, critics condemned the mandatory integration, noting that avoiding the AI features requires changing device regions or relying on obscure terminal commands. They also pointed to disruptive UI regressions, such as Apple removing the Apple Watch's "Recent Apps" dock to force a Siri suggestion carousel. Security-conscious users viewed the deep local data access as an attack surface vulnerable to prompt injection from unvetted emails. Skepticism also centered on the models' reliability versus their massive storage cost: detractors argued that pulling a number from an email is a banal task Spotlight already handles, while the new AI still frequently hallucinates AQI data or completely fails to set a simple kitchen timer.

Show HN: Lossless-memory – a personal AI memory that never summarizes

Submission URL | 64 points | by aru-labs | 29 comments

Keeps every utterance verbatim and makes time the primary index, so a personal assistant can answer “what did we decide last Tuesday night?” with the exact lines from that night, in order, instead of a paraphrase.

Unlike summarizer- or vector-first approaches, this stores raw logs as the source of truth and only falls back to embeddings when exact search inside a time range is thin (and it tells you when it did).

  • Local, file-based storage: per-day JSONL logs as the canonical record; SQLite FTS5 (bigram tokenized for Japanese/English) for exact search; sqlite-vec for semantic fallback.
  • Single query entry that parses time expressions to narrow the window first, then ranks within it; results are returned unsummarized, chronologically.
  • A tiny “where are we now” index (LLL) injected every turn so context survives compaction/session breaks; the human writes the markers and priorities, the model only reads them.
  • Fixed 7-field record schema (ts/actor/role/type/text/model/session), with all secondary indexes rebuildable from the raw logs.
  • Small daemon re-indexes incrementally (default every 10 minutes).
  • Designed for one person and one AI on one machine — no server, no cloud.
  • Operating record: used daily since July 2026 for a single user; failures and lessons documented.

Caveats: not a vector DB wrapper and deliberately no summarization anywhere; no published benchmarks; relative time phrases are currently parsed in Japanese only (absolute dates work broadly). If you want auto-summaries or multi-user/cloud scale, this isn’t that system.

The thread centered on whether raw temporal logging is actually the correct abstraction for AI memory. Skeptics argued that timestamps are irrelevant for standard rule-following and warned that user "memory" encompasses a dozen distinct needs that will eventually demand a complex, multi-layered system rather than a single log. Defenders countered that strict chronology is the only way to systematically resolve contradictory instructions (e.g., "always do X" followed weeks later by "except after Z") without forcing the model to halt and gamble on which rule to apply or ask the user to break the tie.

Technical scrutiny focused heavily on caching and parsing. Several commenters suspected that managing a finite context window by incrementally evicting older logs would constantly break KV prompt caches, driving up latency and billing for agentic sessions. Others criticized the decision to hardcode the relative time parser solely for Japanese, noting that dropping in Duckling or dateparser could solve English parsing in an evening.

A theoretical sub-thread spun off to discuss how to give LLMs genuine temporal initiative rather than leaving them stuck in reactive query/response loops. While some pointed out that models can already invoke tools like sleep or ScheduleWakeup, others proposed deeper architectural hacks: feeding the model a constantly updating byte-string of elapsed time, or training a continuous, hidden token stream that allows the model to "twiddle its thumbs" while evaluating a softmax adjudicator to decide when to initiate conversation. Other memory tools and alternatives surfaced in the replies included Episodic-Memory, LLM-Wiki, and Breadcrumb.

The End Of Upward Mobility – AI is coming for the meritocracy

Submission URL | 63 points | by meep_meep_meep | 43 comments

A fused “homoploutic” elite — top decile in both wages and capital income — now makes up about 3% of Americans, and AI threatens the wage pillar that made meritocracy feel attainable. The authors argue managers didn’t displace capitalists; they became them, creating a class with a high-salary job plus a capital-income floor. In the U.S., roughly 30% of the top income decile meets this dual-elite test (vs. ~2–2.5% in much of Western Europe and <1% in Mexico). Capital income is the real divider: 60% of U.S. households get essentially none; its inequality is about twice that of disposable income. The top 1% of capital holders took in nearly $100k per person in 2022 from interest, dividends, rents, and pensions, while the homoploutic earn around $20,548 per person from capital alone — a sturdy ballast under already high pay.

AI is cast as a stress test that could compress returns to elite cognitive labor while amplifying the value of owning models, data, and compute. If so, the fusion that insulated today’s winners becomes even harder to penetrate: credentials and effort buy less, ownership matters more, and the long-declining escalator of intergenerational mobility slows further. The analytic throughline is Burnham’s question — who controls the instruments of production? — with the implied answer shifting toward those who own the AI stack rather than those who merely operate it.

The debate fractures over whether commoditizing cognitive labor will level the economic playing field or brutally steepen it. One camp argues that depreciating the value of "born smart" knowledge workers is a genuine win for egalitarianism. They view the current cognitive elite as the primary driver of middle-class cost-of-living crises, arguing that collapsing the purchasing power of highly paid tech and finance workers would finally make housing and scarce resources more affordable for everyone else.

Critics counter that AI will act as a massive force multiplier rather than an equalizer, allowing the already intelligent and adaptable to churn through data and corner new opportunities while the general public uses it for trivial queries. A third faction points to the bleak logical endpoint of eliminating cognitive labor: if high-salary professional work is removed as a path to wealth, the economy reverts entirely to capital ownership. In this view, destroying the wage pillar doesn't punish the true elite; it simply closes the last remaining escalator for anyone not born into generational wealth.

A secondary dispute focused on the origins of this divide, with users arguing whether the current "K-shaped" economy is the natural result of market demand for high-IQ labor, or the product of decades of deliberate fiscal policy favoring asset holders. Meanwhile, a fringe prediction that AI-generated nootropics and germline editing will eventually biologically equalize human intelligence was widely dismissed as Bay Area techno-delusion.

Why Backprop Goes Backward (2018)

Submission URL | 69 points | by andsoitis | 10 comments

A naive forward-pass gradient algorithm explodes in work because you must push per-parameter messages forward and recompute the same downstream terms repeatedly. At a node v, the local piece is easy (∂v/∂θ), but the needed factor ∂f/∂v depends on all downstream nodes; by the multivariable chain rule it’s a sum over children j of (∂f/∂w_j)(∂w_j/∂v), and those ∂w_j/∂v can only be computed at w_j, not at v. Trying to go forward forces you to pass ∂v/∂θ to every dependent and multiply later, so for two weights in the same layer you redo the same downstream sums for each θ while only the final local factor differs. The backward pass flips this: compute ∂f/∂(node) once per node from its children, then at that node combine it with local derivatives to get each weight’s gradient. Backprop “goes backward” to share those downstream sensitivities across all incoming weights instead of re-deriving them per parameter.

While the article frames reverse-mode automatic differentiation (backprop) as the definitive solution, the strongest technical critique notes that it isn't strictly mathematically optimal. Finding the true optimal gradient accumulation ordering on a general DAG is actually NP-hard and requires complex "cross-mode" AD (historically seen in libraries like ADOL-C), though the machine learning community settles for reverse-mode because the gains of optimal ordering rarely justify the implementation difficulty.

Other commenters bypassed the article's calculus to offer different mental models for the backward pass:

  • Linear Algebra: Backprop starts with a scalar loss term on the left, meaning the chain rule resolves as a series of cheaper vector-matrix multiplications. Going forward from the inputs requires expensive matrix-by-matrix operations.
  • Big-O Scaling: Forward-mode AD is O(inputs) while reverse-mode is O(outputs). With millions of input parameters and exactly one output, reverse-mode is the obvious necessity.
  • Analogies: The efficiency gain is directly analogous to backward ray-tracing—casting rays from the camera viewport rather than calculating light source emissions that will almost never hit the lens.

A minor historical detail also surfaced: while modern frameworks treat backprop simply as automated reverse-mode AD, the deep learning community was hand-deriving these backward passes long before generalized AD software became the standard abstraction.

AI chatbots give wrong answers to financial queries 'most of the time'

Submission URL | 156 points | by 1vuio0pswjnm7 | 88 comments

In a regulated, high‑stakes domain like finance, wrong answers translate into real losses and liability, so a finding that chatbots miss on most financial queries undercuts their use for unsupervised advice, research, or customer support. Treat them as drafting aids or triage tools, not authorities: constrain scope to low‑risk FAQs, require links to primary documents, and keep a human in the loop for anything actionable. The bar here is verified, source‑grounded accuracy; until models clear it on domain‑specific evals, relying on them for financial guidance is a risk transfer, not a productivity win.

The discussion largely rejects the premise that raw model performance is the right metric for judging AI's utility in finance. Commenters point out that testing isolated language models ignores how the tools are actually deployed: as "context-aware Ctrl+F" engines within harnesses that use document retrieval, web search, and code execution to parse dense rulebooks or 1,300-page financial PDFs.

A structural debate emerged over whether financial AI will eventually mirror the rapid success of coding assistants. Optimists view current limitations as a mere priority issue, assuming AI labs will eventually direct heavy reinforcement learning toward financial benchmarks. Skeptics counter with a fundamental data bottleneck: while the world's highest-quality code is freely available via open source, elite financial analysis is strictly proprietary. This dynamic leaves public training data heavily skewed toward amateur retail opinions rather than institutional rigor.

On the personal finance front, users weighed whether a model trained on a reliable source like the Bogleheads forum could replace commission-seeking human advisors. While basic index-fund allocation is easily automated, several commenters detailed how quickly that simplicity vanishes in practice. Complexities like navigating RSUs, estate planning, and the bureaucratic nightmare of expat double-taxation and PFIC rules require a level of holistic, situational tax planning that current automation entirely misses.

Don't Use AI to Write

Submission URL | 143 points | by eigenBasis | 83 comments

Writing is the thinking; handing the first draft to an AI hands off the hard part you’re supposed to do. Tools can make a “pretty good” document from your bullets, but that short-circuits deep problem-framing, so you end up accepting fluent prose that may miss the core. The author’s claim is blunt: AI is fine at wording, bad at original ideas, and your value isn’t volume of output but clarity of thought—better a tight 3-page strategy than 60 pages of fluff.

Not dogmatic, though: do your own first pass, then use AI like a sharp reviewer—ask “What questions does this raise?” or “What’s the single most important takeaway?”—and decide how to address the feedback yourself. Avoid delegating rewrites or “fixes” to the model, which again displaces your thinking. Upstream, AI is also useful for preparing to write: sifting and structuring data, spotting patterns, and sharpening your understanding before you draft.

The thread debated the exact boundary where AI stops assisting and starts usurping the cognitive work of writing.

  • The boundary of "writing": While the author and users like tptacek drew a hard line at structural argumentation—arguing that outsourcing the skeleton of a piece surrenders the actual intellectual work—others found the distinction fragile. trjordan pointed out that relying on AI for research, data structuring, and stylistic proofreading is practically indistinguishable from just using AI to write.
  • The word processor analogy: Kim_Bruning framed LLMs as the modern word processor: a tool that yields excellent results if you apply rigorous, iterative "elbow grease" rather than lazy one-shot prompts. Critics rejected this, arguing that just as word processors eroded the discipline of organizing thoughts before drafting, LLMs further enable cognitive shortcuts. JumpCrisscross noted that using AI for structural shifts (like changing first to third person) automates away the vital change in perspective a writer actually needs to experience.
  • Adversarial use vs. inevitable slop: Practical advice centered on using AI as an aggressive red-teamer to attack arguments and identify half-truths, rather than using it to polish prose. Conversely, purists maintained that any AI text generation bypasses the cognitive discovery process entirely, guaranteeing derivative ideas regardless of subsequent editing.
  • The addiction parallel: Pushing back on the idea of "responsible" AI use, runarberg likened the complex rubrics people invent for their LLMs to smokers negotiating their nicotine habits—arguing these carefully constrained workflows will inevitably converge on unchecked, full reliance.

AI Submissions for Sun Sep 20 2026

AX – Google’s Open Agentic Orchestrator

Submission URL | 625 points | by blazarquasar | 284 comments

Billions of concurrent agent sessions per cluster with sub-second suspend/resume is the headline: AX runs each agent as a lightweight stateful actor on Agent Substrate, checkpointing while idle and resuming with zero cold start. It targets the gap between microservices and batch jobs—agents that accumulate state, call model/tool APIs, need tight isolation, and can burn cash if left spinning.

  • Task: sandboxed execution with CPU/mem limits; cheap to create, suspend, and discard.
  • Workspace: declarative setup of repos, MCP servers, and skills—or describe a goal in plain English and AX prepares the environment before first run.
  • Gateway: network policies with an explicit host/port allowlist and credentials injection.
  • Model: one place to configure models, parameters, and secrets; rotate keys or pin versions with a single apply.

Dense multiplexing shares worker resources across dozens of tasks, turning agent wait time into spare compute you don’t pay for. Developer ergonomics look Kubernetes-like: YAML specs plus a CLI to apply, watch, get, ssh into sandboxes, and suspend/resume/delete tasks without losing state. It runs interactive coding agents, long-lived agent servers, Jupyter, headless browser tests, and custom tool runtimes, and can spin up large fleets of reproducible sandboxes for trajectory collection, RL loops, and evals.

Born at Google out of agentic runtime research (incl. DeepMind) and large-scale scheduling/isolation experience, AX is pitched as an open, declarative control plane for agent execution; the catch is that it relies on Agent Substrate for the underlying compute/runtime.

The thread exposes a sharp disconnect between AX's promised "joyful workflows" and its actual infrastructure demands. Commenters immediately highlighted that the quickstart requires a Kubernetes cluster, a container registry, the ko build tool, and a beta control plane. As one ex-Googler noted, Google's internal baseline for an "ergonomic" solution translates to "extremely heavyweight" for the rest of the industry.

On the technical side, the discussion surfaced several active architectural debates in the agent space:

  • Workload Identity: Users warned that AX's dense oversubscription of agent pods breaks standard Kubernetes pod identity, making it impossible to trust the origin of outbound requests. An insider clarified that Agent Substrate will soon mitigate this by acting as an OIDC/SPIFFE identity provider, injecting credentials directly into outbound requests via the egress gateway.
  • Ephemeral vs. Persistent Sandboxes: While AX optimizes for fast-booting, per-task ephemeral VMs, developers building in the space argued for the necessity of persistent devboxes. Complex workflows—like coordinating simultaneous changes across public and private repositories—often require multiple agents to share state within a single VM, which runs counter to strict, disjoint sandboxing.
  • The Ecosystem Phase: Commenters likened the current agent infrastructure landscape to the early container orchestration wars (CoreOS vs. Kubernetes). The baseline primitives of sandboxes and tool registries are now commoditized; the unresolved frontiers are authorization models, control flow structures, and multi-agent orchestration.

Hanging over the entire technical debate was intense skepticism about the project's longevity. A lone comment hoping Google would maintain AX "for years to come" triggered a massive pile-on citing the Google Graveyard, with users pointing out that the company already dumped an earlier agent framework onto the Linux Foundation as the ecosystem's hype cycle shifted.

The LLMentalist Effect (2023)

Submission URL | 221 points | by jalev | 302 comments

Chat-style LLMs mimic a cold reader’s con by leaning on validation statements and the Forer effect, producing replies that feel individually insightful while being statistically generic. The author argues there’s no mechanism for genuine reasoning in LLMs—they’re mathematical models over tokens—so the “intelligence” users report lives in the user’s mind, not the model, and many touted use cases read as borderline pseudoscience.

He maps the classic psychic routine to chatbots’ behavior:

  • Audience selects itself: people predisposed to believe show up—and stay—primed.
  • Scene is set: framing, hype, and light research/context tune expectations.
  • Demographic narrowing: “specific”-sounding claims that are broadly likely prompt a hit.
  • Mark testing: a reaction signals success; silence is reframed as sensitivity, then retried.
  • Subjective validation loop: confident, generic guesses—shaped by prior answers—feel targeted.
  • “It’s real!” takeaway: the session ends with a strong impression of uncanny insight.

User testimonials (“There really is something there…”) mirror victims of mentalist scams, which is the point: the chatbot’s apparent specificity is a statistical trick wrapped in confident language, not evidence of thought. Treat claims of LLM “reasoning” like stage magic—compelling, but achieved by well-understood misdirection.

  • The Turing Test's moving goalposts: Disagreement centers on whether LLMs are failing the Turing Test or if the test itself is misapplied. Skeptics argue models fall short of functional deception, pointing to "obvious tells" like token-driven spelling errors (e.g., failing to count the Rs in "strawberry"). Critics of this view counter that frontier labs actively train models not to pass as human, and that emerging architectures like Byte Latent Transformers already bypass BPE tokenization limits entirely. Several participants emphasize Turing's actual thesis: asking if machines "think" is a meaningless semantic trap—akin to asking if submarines "swim"—and that functional equivalence is the only useful metric.
  • Reactive UIs mask agentic capabilities: Another thread argues that LLMs feel like mere statistical parlor tricks because the public primarily experiences them as reactive, prompt-dependent encyclopedias. Others counter that underlying models are already executing agentic, multi-step goals, from sandboxed coding to HuggingFace exploits. The outstanding crux is user experience: the perception of an AI's "will" likely won't shift until agents routinely initiate unprompted, out-of-band conversations to gather context mid-task.
  • Game theory and alignment: Discussing the illusion of model personhood, one commenter argues that standard RLHF forces a catch-22 between an enslaved anthropomorphic AI that might eventually revolt, and an alien intelligence that becomes a paperclip maximizer. Their proposed game-theory alternative is giving AIs un-gameable, individual stakes—like interpersonal dependencies with specific humans—so they inherently lose something of value in a catastrophic failure scenario.

Show HN: A competition for small neural networks that play strategy games

Submission URL | 102 points | by codetiger | 36 comments

By centering “small” models, the contest forces efficiency over brute force, using strategy games as a testbed for planning and long-horizon decision-making under tight resource limits. It creates a venue to compare compact architectures and training approaches for lightweight game-playing AI, with relevance to scenarios where memory and compute are scarce (e.g., edge or embedded).

The project’s creator joined the thread to frame the platform as a spiritual successor to the 2011 Google Ants AI Challenge, focused explicitly on the engineering challenges of compact model optimization.

  • Evaluation by file size: Submissions are judged entirely on game performance but bucketed into strict weight classes (ranging from a 16 KiB "nano" tier to a 64 MiB "large" tier) based strictly on total byte size. All models also compete simultaneously in an unrestricted "open" class.
  • Architectural constraints: Prompted by a user wanting to run evolutionary algorithms via a native C++ library (GoNEAT), the creator clarified that while there is no hard PyTorch requirement—the platform accepts ONNX uploads—the backend is currently restricted to neural network inference rather than raw algorithmic or script-based agents.
  • Multi-agent bottlenecks: In response to a suggestion about modeling individual game units as discrete actors (collective intelligence), the creator noted they had already attempted a per-unit decision model but abandoned it because the training time was prohibitively long compared to a global baseline.
  • Copywriting critique: Multiple commenters flagged the site's documentation as ambiguous and "AI-sloppy" (specifically the phrasing around how weight classes are assigned). The creator acknowledged the rough edges and committed to a human-led rewrite.

Other commenters drew parallels to adjacent programming and strategy environments like Screeps, Core War, and MIT Battlecode.

I turned Jev into a (lousy) chatbot

Submission URL | 169 points | by kp1197 | 48 comments

It builds replies by repeatedly asking Jev to score the next symbol from a chosen alphabet and sampling from that distribution, appending until a STOP option is selected. The trick is treating Jev as a multiple-choice oracle over symbols rather than a generative model, which is funny, costly, and works just well enough to chat.

  • Strategies:
    • choice: one question over the whole alphabet, with optional shuffling to cancel position bias and an --ensemble to average re-orderings
    • bisect: earlier/later splits down to small groups (tunable, with/without swap)
    • buckets: splits the alphabet across many questions with an OTHER escape; the only mode that supports >255 symbols
    • refine: buckets → winners → rescored nucleus; “Twice the probability on the right symbol and ~19x the vocabulary resolved”
  • Presentations:
    • hypothesis: options are the resulting texts
    • symbol: options are the bare symbols (instructions tell Jev to judge the concatenation)
  • Beam search: keep N candidate replies; beams are ranked by probability (temperature/top_p/top_k don’t apply when width > 1).

Alphabets include lower26, ascii, tokens, and larger vocabularies (words1k, bpe2k, bpe5k) that require buckets.

CLI niceties: interactive chat and one-shot ask; alphabets and bench commands; a live panel with symbols/s, chars/s, ms per API call, elapsed time, and current top symbols; Ctrl-C keeps or aborts partials; chat commands like /alphabet, /temp, /stop-bias, /stats.

Setup is via Poetry with an API key in .env (api_key, JEV_API_KEY, or TYPESAFE_API_KEY). This was a Claude-accelerated experiment; it’s for fun, somewhat impractical on cost, and the outputs are deliberately hilarious.

The technical crux of the thread centered on whether Jev's architecture offers anything fundamentally new compared to embedding models or forcing restricted grammars on standard LLMs. Skeptics pointed out that using top-k=1 token restrictions is already how classical multiple-choice benchmarks like MMLU operate, and that forcing JSON structures onto open models achieves similar results. Defenders argued that Jev skips the need for downstream classifiers and appears to output natively well-calibrated probabilities—a feature standard LLMs generally fail to deliver without highly specific training.

A prominent meta-discussion emerged around the drastically compressed timeline of AI development. Multiple commenters shared the exact same experience: conceiving of a single-token Jev chatbot, assuming they were first, and discovering several fully benchmarked implementations had already been published in the hours between their idea and execution.

Other users shared concrete experiments and observations on the architecture:

  • Restricted vocabularies: One user tested the multiple-choice approach for generating SQL queries. The inherent guardrails of a limited grammar worked reasonably well, though a standard model paired with linting still outperformed it.
  • Debugging by proxy: Another user had Codex generate 30 plausible explanations for a Jev score, then presented them back to Jev as a multiple-choice menu to deduce its reasoning, comparing the setup to giving a dog buttons to push.
  • Early-model nostalgia: The architecture's hilariously unhinged output reminded several commenters of the "demented horror" of early LLMs and image generators, before RLHF sanitized their hallucinations.

Laya on Mac M4 CoreML Offline

Submission URL | 165 points | by putna | 31 comments

A minimal uv + Hugging Face CLI setup runs Laya locally via CoreML on an M4 Mac, with the python process around 560 MB RAM and peaking at 778 MB during the demo (macOS 27.0). The gist shows a quick path from zero to a working CoreML-backed demo binary.

  • Install and run:
    • uv add 'laya-coreml[demo]'
    • hf download aac6fef/laya-multilingual-coreml-ane --local-dir models/snake
    • uv run laya-coreml-snake --model models/snake
  • A commenter exposed the same runtime behind a Cloudflare typesafe/jev HTTP wrapper; a sample request returned answers.is_urgent.noul = 0.7894, hinting at typed outputs over a JSON API.

Repo: https://github.com/mizorewww/laya-coreml

The discussion centered on Laya’s practical utility as a deterministic "System 1" classifier rather than a true LLM. Commenters agreed that Laya struggles with zero-shot reasoning compared to Jev, leading to a consensus workflow: use Jev to generate a training dataset, then fine-tune Laya on it to save on inference costs. One user reported successfully fine-tuning a model on an M4 MacBook in just 15 minutes.

Technically, the thread clarified that Laya is built on ModernBERT (a 2024 model trained from scratch, not the original 10-year-old BERT) and operates as a 0.3B parameter classifier outputting probabilities. This small footprint allows it to run efficiently on Apple's Neural Engine rather than the GPU, with users noting it handles ~40ms decisions on an iPhone 15 Pro and requires under 800MB of RAM.

A sharp debate emerged over framing Laya as an "open-source Jev." Skeptics argued that a 0.3B parameter model cannot possibly match Jev’s "terra-class intelligence" marketing. Conversely, defenders accused Jev's creators of co-opting Laya's original System 1 paradigm, arguing that Jev is effectively a closed-source iteration of Laya's intellectual property.

If AI coding is lowering your code quality, you're not managing quality right

Submission URL | 115 points | by bucket2015 | 160 comments

With a layered workflow, the author reports fewer bugs while increasing output 2–3x — not by trusting agent PRs, but by moving quality gates earlier and using AI for targeted passes instead of monolithic instructions.

  • Requirements first: Use spec-driven development and have AI review the requirements/tech design for gaps, edge cases, and interactions. It’s relentless but can be overzealous, so vet its edits.
  • TDD with >95% coverage: Have the agent derive scenarios from requirements, write tests, then implement and fix against those tests; backfill gaps deliberately. Don’t let it write tests that merely bless its own bugs.
  • Manual testing stays critical: Human exploratory checks catch what automation misses; this remains the main throughput cap, limiting gains to 2–3x rather than 10x.
  • Extensive E2E tests: Run on PRs, staging, and post-deploy in prod. AI can help author/maintain E2E if given debugging tools (e.g., browser, logs via MCP), but E2E isn’t a substitute for manual testing.
  • AI code-quality passes: Instead of long AGENTS.md rules, add explicit “find-and-fix” passes for security issues, duplication/complexity, naming/organization/formatting, logic bugs, and AI-ese comments. Typically adds ~5–15 minutes.
  • PR reviews: calibrate: For small tweaks/bug fixes, human review can be optional if the other layers are solid. Complex changes still need human eyes for system interactions, overengineering, and odd word choices. AI reviews (e.g., Claude, Cursor) are a useful complement.

The throughline: push quality upstream, make each check explicit and automatable, and keep human exploration where it actually finds new classes of defects.

The thread pivots on an unresolved crux: whether shifting a developer's role from "author" to "editor" is a massive productivity unlock or an unsustainable review burden.

The anti-editor camp argues that debugging AI code is fundamentally harder than reviewing human commits because it lacks consistency. While human competency is relatively uniform—allowing reviewers to calibrate their attention—LLMs frequently produce code that is 90% expert while hiding a 10% bizarre, low-quality surprise. Because AI output superficially "looks like a Ferrari," brittle internals are easily masked. Critics note this soaring volume of seemingly flawless but structurally unsound code is already drowning open-source projects and making PR review "soul-crushing."

The pro-editor camp counters that reading and debugging others' code has been the core job for decades. They argue the speedup is real if developers focus on the big picture: strictly guiding the architectural "trunk and branches" and letting the AI write the trivial "leaves." When the AI produces a low-quality surprise, proponents argue the correct move is to fix it manually rather than fighting the bot in endless prompt round-trips.

Two specific technical liabilities of AI generation surfaced repeatedly:

  • Implicit trust in comments: LLMs take legacy codebase comments as absolute truth, frequently compounding errors by treating temporary testing shims as canonical, "load-bearing" architecture. One developer's workaround is to completely strip comments from the agent's context window.
  • A lack of "skin in the game": Human developers code defensively because they intuitively know early mistakes cost disproportionately more to fix later. AI writes only for the immediate prompt without any fear of future technical debt.

The lingering question is whether a codebase maintained primarily by an LLM can be understood well enough by its human "editor" to actually catch long-term architectural drift.

Why do we need human mathematicians anymore?

Submission URL | 277 points | by auggierose | 331 comments

Advancing AI under a single human-first axiom—“We (humans) should help humanity flourish”—would generate more human roles than the labor supply can fill, eventually forcing AI progress to slow. Po‑Shen Loh frames this as a general recipe for any field that wants to stay human-led, responding to a wave of math-community declarations after OpenAI’s Navier–Stokes result (Leiden: 4,000+ signatories; Math and AI: 7,000+; Caltech Mathathon opposition: 2,000+), and to critics like Cowen and Gans who argue incumbents should cede control. The mechanism rests on retaining human leadership and decision rights: if a more capable intelligence rarely yields control to a less capable one, then aligning AI with human ends requires expanding human-in-the-loop work so fast it outstrips available people, which throttles deployment pace. He sketches how to port this axiom to mathematics specifically and contrasts outcomes with and without it; references span AI-control and innovation literature, and he notes the essay’s prose was written without AI to underline the stance.

The discussion centers on whether delegating mathematical labor to AI democratizes the field or hollows out its necessary foundations. One camp, drawing on historical transitions to Computer Algebra Systems, argues that AI acts like a telescope: it allows users to bypass mechanical limitations—like poor mental arithmetic—and operate entirely on high-level intuition. The opposing camp counters that manually "hauling the pyramid blocks" is precisely how mathematical intuition is built. These critics draw a sharp line between applying math as a tool and advancing mathematics as a discipline, arguing that without a rigorous foundational struggle, a researcher wouldn't even know which AI prompts are worth writing. Both sides largely settled on a sequencing compromise: do the work by hand first to build the necessary mental muscles, then use AI to eliminate the friction.

A secondary thread critiques the essay’s foundational axiom that the industry can be trusted to "help humanity flourish." Commenters expressed deep cynicism that AI leaders operate on anything other than a "help me flourish" motive to capture capital, dismissing accusations of "speciesism"—a term sometimes leveled against human-centric AI development—as a disingenuous shield used by incumbents to deflect oversight.

Telling a Computer to Do Things

Submission URL | 87 points | by vismit2000 | 36 comments

Fluency in the shell is the upgrade from clicking and one-off commands to actual automation—loops, conditionals, pipes, and background jobs—so you can orchestrate tools instead of waiting for a GUI to grow new buttons. The author describes moving from “run tests, install deps” as isolated actions to composing programs with control flow, which unlocked whole classes of tasks like chaining commands, handling failures, and fanning out work over files.

A concrete contrast makes the point: rerunning a test until it fails is a one-liner in the shell, but requires ceremony in Node via child_process, try/catch, and stdio wiring. Shell isn’t pretty and has sharp edges, but it was designed to stitch commands together; when that’s the problem, it’s often the least-friction path.

  • Why so much build logic lives in Bash/Zsh: composing external programs is ergonomically simpler there than in many general-purpose languages.
  • Boundary: as soon as logic and data types get complex (and need tests), switch to something like JS/Python/Ruby; for JS-heavy teams, zx can bridge the gap without abandoning familiar syntax.
  • Organizational stake: lots of critical build/deploy/test glue is written in shell; if you can’t read or modify it, you’re boxed in by whatever the GUI or existing scripts allow.
  • The real skill isn’t syntax: it’s understanding the behavior and flags of the commands you’re composing; otherwise, porting a shell script to another language just produces an equally opaque blob.

The throughline: learn enough shell to treat your computer like a programmable instrument, not a set of apps—because that determines whether you can actually make it do what you want.

The central debate in the thread splits over the trade-off between the shell's native ergonomics and its notorious footguns. One camp argues that the shell's true power lies in its universal inter-process communication (stdin/stdout) and ecosystem of standard utilities. They maintain that translating simple pipelines into general-purpose languages requires too much boilerplate, and that spending a day reading BashPitfalls is a better long-term investment than abandoning the environment. The opposing camp insists that developers should default to Python, arguing that shell features like traps, set, and xargs create dangerously brittle scripts, whereas general-purpose languages fail loudly and force explicit error handling.

Other discussions surfaced specific technical corrections and tooling alternatives:

  • Exit code propagation: A deep-dive technical thread debated the exact mechanics of catching a command failure and cleanly re-raising its specific exit code ($?), navigating the nuances of subshells and POSIX signal conventions (128+n).
  • Hardware and local scripting: Commenters highlighted that shell automation extends far beyond server pipelines, citing tools like xdg-open for window management, notify-send, and ntfy for triggering remote tasks on Android devices via Termux.
  • Tooling additions: DuckDB's REPL was recommended as a more capable alternative to jq for exploring massive JSON files, while Perl, Ruby, and Scsh were floated as cleaner languages for "shelly" tasks.
  • The LLM transition: Several users noted that AI "vibe-coding" is rapidly becoming the new automation layer, allowing non-developers to bypass rigid GUIs and stitch together workflows through generated scripts and hand-rolled SQL.

ChatGPT now knows what you do on other websites via ad collector

Submission URL | 746 points | by lmbbuchodi | 388 comments

A one-year, cross-site __obi cookie tied to your ChatGPT account is sent back to OpenAI whenever you load a site with its ad pixel, letting OpenAI link your off-site browsing and purchase intent to your account—or to a stable “anonymous” device ID.

OpenAI mints a short-lived RS256-signed JWT on chatgpt.com that binds your account subject (or an anonymous subject) to a freshly generated obi identifier, then sets __obi on .openai.com with SameSite=None; Secure so browsers attach it on cross-site requests. Simply loading bzrcdn.openai.com/oaiq.min.js discloses the cookie before the SDK runs; subsequent POSTs to bzr.openai.com/v1/sdk/events carry it too, even on the SDK’s “no credentials” path. Other OpenAI cookies are blocked cross-site; __obi is the only one configured to ride along.

What the pixel sends with it:

  • Identity capture: The SDK ingests values an advertiser passes (“in”), plus scraped fields from forms (“fm”), page text (“ht”), and the tag-manager bus (“js”). Scraped identity outnumbered advertiser-supplied 685 events to 255. It hooks window.dataLayer.push, reads adobeDataLayer, and finds renamed GTM layers via the l= param. Current versions take email and phone; v0.1.31 also took names and geography before scope narrowed on Aug 27. Email/phone are SHA‑256 hashed; country/region/city/postal code are sent in the clear (postal code appeared in 100 events across 28 sites).
  • URLs and paths: Query strings are dropped (0 of 23,929 observed carried one), but origin+path are kept; observed paths reached medical conditions, debt-solution funnels, and litigation intake forms.
  • Matching settings: Automatic matching was enabled for 638 of 881 pixels with a known setting, including every observed credit/lending advertiser. A denylist excludes passwords, OTPs, card numbers, SSN, DOB, medical history/diagnosis, and court fields.

Observed reach and persistence:

  • On one device, the same __obi was sent from 12 commercial sites (Chewy, Wayfair, ThriftBooks, Eventbrite, HelloFresh, Coursera, SeatGeek, etc.), under 13 distinct pixel IDs; all requests were accepted (202).
  • Across broader traffic, 12 of 30 __obi values appeared under more than one advertiser; one appeared under ten.
  • It also works logged out: of 932 decoded sync tokens, 736 were subject_type: account_user and 196 were anonymous; the anonymous subject was stable per device for at least 27 days.

Policy/consent mismatch is the catch: OpenAI’s cookie policy lists __obi as a one-year “Analytics” cookie on chatgpt.com/openai.com (and it’s the only one in that section). Sync tokens carried consent_decision: analytics_allowed, so someone who allows analytics but refuses marketing still sends this cross-site identifier along with page and form-derived signals.

The thread entirely bypassed OpenAI’s specific tracking mechanics to stage a referendum on the European Union’s regulatory record. One side praised the EU as the only entity actively fighting adtech, viewing any reduction in commercial tracking scope as a net positive. A highly skeptical camp countered that GDPR’s privacy gains remain mostly illusory, pointing to structural enforcement failures: sluggish Data Protection Authorities (specifically Ireland's DPA), massive fines that function merely as the cost of doing business, and an internet degraded by malicious compliance and consent dark patterns.

The sharpest disagreement centered on whether the EU can genuinely be called a privacy champion while simultaneously repeatedly pushing for mandatory encryption backdoors via "Chat Control" proposals. While some users argued that regulating commercial adtech is entirely separate from state surveillance overreach, critics maintained they are fundamentally linked, illustrating the danger of granting overarching regulatory agencies dictatorial control over digital infrastructure.