AI Submissions for Tue Sep 29 2026
Livenerf: Has Opus 5.5 been nerfed yet?
Submission URL | 862 points | by bryan0 | 369 comments
Livenerf is building a launch-week baseline before anyone can argue from memory: it runs a fixed panel of 78 questions daily and compares later 10-day windows against that baseline, with prompts, graders, CLI version, and raw logs pinned. It uses headless Claude Code on a Claude Max subscription rather than an API key; output-token counts are tracked because reduced effort may show up there before accuracy shifts.
As of September 30, seven of the planned 30 days had completed, with 90 samples per day and no missed runs. The benchmark estimates it can detect an accuracy change of about 7.5 percentage points per 10-day window, but validation could not distinguish Opus 5 from Opus 5.5—so a smaller or same-family serving change may slip through. No regression result exists yet; the first comparison is planned after the baseline and next 10-day window.
Much of the discussion centers on whether perceived model degradation is an engineering reality or a psychological illusion.
Several commenters argue that "nerf" complaints are largely driven by human perception:
- Expectation drift: A former chat-app developer noted that even when operating a completely frozen stack with zero server-side changes, users regularly complained models had been degraded. As users grow accustomed to a tool, expectations rise while prompting discipline grows lazier, causing perceived performance (
actual performance / expectations) to drop. - The illusion of randomness: Humans naturally misinterpret normal stochastic clustering as systematic degradation, reading intentional changes into random runs of poor outputs.
Conversely, others argued that silent quality degradation is a pervasive, rational corporate playbook. Commenters pointed to manufacturing cost-cutting and "shrinkflation"—such as the gradual weakening of IKEA furniture materials or Amazon Basics allegedly using premium OEMs (like Eneloop batteries or Corning fiber) to build early reviews before swapping in cheaper suppliers.
The consensus landed in the middle: while user hysteria and honeymoon falloff generate constant false alarms, real regressions do happen. Commenters noted that independent tracking tools like Nerf Bench previously caught an actual regression in Opus 4.6 that Anthropic later publicly confirmed, underscoring the need for empirical monitoring rather than relying on user sentiment.
Language models for text classification: From bag-of-words to Jev
Submission URL | 197 points | by Anon84 | 10 comments
Jev sits between general-purpose LLMs and task-specific classifiers: it trades some breadth for faster, cheaper classification, while avoiding the setup of a model built for one narrow task. Raschka puts that pitch in historical context, tracing text classification from bag-of-words features and classic models through neural networks and transformers. He says his view shifted from “I can build this myself” to being impressed by how well Jev works, but the article’s account of its methodology is explicitly an educated guess.
Commenters focused on cutting through the launch hype to evaluate where the model actually fits in production:
- Calibration is the real operational unlock: While general interest focused on raw accuracy, practitioners pointed out that reliable confidence scores matter far more for production routing (e.g., auto-handling inputs over a 0.9 confidence threshold and sending the rest to human review). Plain fine-tuned models often overfit negative log-likelihood and produce wildly overconfident probabilities; training confidence directly addresses this. However, benchmark skepticism remains: without knowing whether standard test sets like IMDb were included in the synthetic training mix, published accuracy figures should be treated as ceilings rather than reliable baselines.
- Misleading game demos obscured the architecture: The launch suffered from familiar hype distortion, with onlookers claiming leaps toward AGI or instant visual processing. Showcases involving Doom or Minecraft led many to believe the model possessed ultra-fast computer vision, when it was actually operating on parsed game-state text.
- Bridging the architectural history: Readers appreciated tracing the lineage from bag-of-words up to modern architectures, though one noted a missing pedagogical link: continuous bag-of-words (CBOW) and embedding-averaging approaches (like fastText), which first bridged discrete vocabulary counts and neural semantic spaces.
GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Submission URL | 1041 points | by crorella | 926 comments
GPT 6.1 Sol is pitched as approaching Astra’s intelligence at one-fifth the price. The title doesn’t specify what the price comparison covers or how the intelligence gap was measured.
While benchmarks cited in the thread place GPT 6.1 Sol comfortably ahead of Opus 5.5 and Sonnet 5.5 in coding environments at a fraction of the API cost, the discussion quickly turned to whether those savings hold up in practice.
The core tension centers on Codex’s context management. Commenters praising the model highlighted its 50% cheaper cached input ($0.10 per million tokens) and reported running dozens of PRs and investigations for mere dollars. However, multiple developers pushing the tool through heavy workloads reported exhausting subscriptions rapidly due to aggressive context compaction every few minutes. Where some users run massive, unattended 400k–700k token sessions on Claude without issue, Codex’s default ~256k–275k window forces compaction cycles that can degrade working memory and blow through cache hits.
The divide split the thread into two camps on workflow:
- The orchestration argument: Several developers argued that compacting every few turns is a workflow mistake rather than a model defect. They advocate breaking work into bounded tasks, using tools like linear tickets, relying on OMP’s checkpoint/rewind feature, or delegating to fresh subagents to scale context linearly rather than ballooning a single monolithic prompt.
- The developer ergonomics critique: Others countered that Claude’s large native context remains vastly more practical for deep repository work, avoiding the drift that occurs when Codex compacts and loses grasp of edge cases. For those hitting the wall, participants pointed out that Codex’s compaction threshold isn’t hardcoded; it can be manually raised to ~700k via
~/.codex/config.toml, though OpenAI leaves the setting largely undocumented.
On model behavior itself, engineers noted a stylistic difference: 6.1 Sol demonstrates significantly lower initiative outside explicit instructions compared to Anthropic’s models—a trait several commenters welcomed as a relief from agents hallucinating unrequested features.
Dots: Always-on agents
Submission URL | 744 points | by alvis | 627 comments
Dots is positioned around agents that stay active rather than waiting for a prompt. The title gives no details on what they do, what “always-on” means in practice, or how users control them.
The discussion largely turned into an architectural show-and-tell on how to run autonomous, "always-on" agents in practice, paired with a sharp debate over whether unattended AI can ever be trusted.
Practitioners running multi-agent setups shared surprisingly low-tech, resilient patterns:
- UNIX sandboxing over complex frameworks: Rather than reaching for Docker or MCP, operators reported running agents inside dedicated UNIX user accounts, relying on standard OS permissions to restrict access to files and tools. Inter-agent coordination and human check-ins are routed via local Maildir and plain email, with systemd timers and POSIX locks waking agents up to tackle background tasks—such as triaging bug backlogs, clearing disk space on pet servers, or prepping support replies—during scheduled idle time.
- Continuous context over fresh sessions: Instead of spinning up clean context windows per action, several users keep permanently rolling sessions alive by leaning on model-native compaction, supplemented simply by having the agent maintain a rolling diary and wiki in its home directory.
- The latency and cost penalty: Commenters noted that always-on workflows can multiply inference spend by two to three times and drastically increase wall-clock completion time (e.g., an hour asynchronously versus ten minutes of synchronous prompting), but argue the trade-off is worth it to eliminate active human babysitting.
The pushback centered on risk, with skeptics calling hands-off execution reckless. In response, operators argued that managing an agent is closer to leading an easily distracted junior employee than running deterministic software: neither is faultless, both require clear guardrails, and trust is granted incrementally as task accuracy is proven. Others advocated for cross-model auditing—pitting rival frontier models against each other to vet decisions before execution—though critics cautioned that the sheer volume of autonomous output inevitably tires humans into rubber-stamping changes they never actually verified.
A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf]
Submission URL | 421 points | by damaru2 | 136 comments
The paper examines privacy in web- and mobile-based conversational AI agents, but the available text contains no findings or methods to summarize more specifically.
Commenters turned the discussion toward specific, under-the-radar privacy leaks across commercial chat interfaces:
- Unfinished prompt exfiltration: Users noted that ChatGPT's web client periodically posts unsent, in-progress drafts to a
conversation/prepareendpoint. While some suggested benign explanations like cross-device draft syncing, cache pre-warming, or human-typing cadence verification to detect API scrapers, others argued it creates an end-run around privacy policies: text entered into an input box can be logged, profiled, or ingested long before a user actually hits "submit." - Security-by-obscurity in chat URLs: Several commenters criticized services like Perplexity and Gemini for relying on UUID-based links without explicit authentication to gate access to chat logs. While mathematically unguessable, these URLs routinely leak into cleartext local browser histories, malicious extensions, aggressive web prefetchers, and public search indexers—recalling an incident where Claude artifacts were scraped and indexed en masse by search engines.
- The case for local models: Referencing recent controversies where researchers' draft mathematics in private Codex sessions were reportedly ingested into model improvements, commenters argued that terms-of-service hair-splitting ("user prompts" vs. "telemetry scratchpads") makes true privacy impossible on hosted platforms. The consensus among technical users was pragmatic: if proprietary data or unformed ideas cannot risk exposure, running local open-weight models is the only architecture that provides a reliable boundary.
Show HN: TurboGPT: train 22KiB transformer in 13s
Submission URL | 53 points | by lostmsu | 10 comments
The 22 KiB byte-level GPT is implemented in CUDA C++ and inspired by minGPT home experiments. The repo reports 2.5295 bits per byte on hn1g after 1.5 billion training tokens; the 13-second headline run doesn’t specify its hardware. It’s MIT-licensed and requires CUDA 13.4; the README documents Nix/Linux and Windows builds, though the author says the build script is Windows-only.
Commenters immediately questioned the purpose of yet another miniature GPT implementation, arguing that the endless stream of these projects adds little pedagogical value beyond Andrej Karpathy’s original minGPT and nanoGPT, suspecting resume-padding or LLM-assisted code churn.
The author (lostmsu) pushed back with a practical hardware justification: while minGPT is great for teaching concepts, it is too slow for iterating on novel transformer ideas at home, and nanoGPT targets multi-GPU datacenter nodes (like 8x A100 setups). A fast, raw CUDA implementation makes rapid, single-machine experimentation feasible on architectures that lack pre-baked, optimized primitives.
A few technical observations rounded out the thread:
- Disk size vs. parameters: One commenter noted the odd trend of headlining models by their file size on disk rather than their parameter count.
- Optimization limits: Another half-joked that with a model so small, one could almost skip standard gradient descent and solve the KKT conditions directly.
- Missing training data: When an initial link to the
hn1g.txtdataset was reported dead, the author uploaded it to Hugging Face, clarifying that the byte-level predictor can run on arbitrary raw text.
Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound
Submission URL | 64 points | by tgluck | 15 comments
The 98% target means matching Jev, not being right: Jevstiller’s local model answers only when it is confident and the request looks in-distribution; otherwise it falls back to Jev. It uses a frozen 384-dimensional sentence encoder and a small logistic-regression head trained on Jev’s probability outputs, answering locally in about 15 ms on CPU.
The router calibrates a bound on the share of all requests where its answer would differ from Jev’s. Instead of choosing a confidence threshold from a point estimate—which exceeded the 2% disagreement budget on 6–12 of 20 splits per task—it tests candidate thresholds with an exact Clopper–Pearson confidence bound. Across five public tasks, the bound-based rule broke the budget once in 100 splits, consistent with its stated 95% guarantee, while giving up some local coverage.
That guarantee is statistical and depends on the calibration setup; it says nothing about correctness against ground truth.
Discussion focused on the scope, implementation details, and practical constraints of distilling Jev calls into local models:
- Scope and utility: Skeptics argued the proxy only helps with simple text classification—the least interesting application of Jev. The author defended the focus, noting that high-volume, repetitive classification is both the explicit design target and representative of many real-world production workloads.
- Architecture details: In response to technical questions, the author clarified that unlike tools such as model2vec (which compress the encoder itself), Jevstiller pairs an off-the-shelf frozen encoder (defaulting to
bge-small) with a lightweight multinomial logistic regression head written in roughly 80 lines of plain NumPy and trained via full-batch Adam. It acts as a drop-in replacement by pointing the API base URL at the proxy. - Handling errors and terms of service: Commenters questioned whether caching and distilling outputs breaches provider Terms of Service, and asked if human corrections could override Jev's mistakes. The author acknowledged that agreement does not equal ground-truth accuracy—coverage drops significantly on noisier tasks—and noted that while manual label overrides are an interesting future direction, the system currently only optimizes for fidelity to Jev's original outputs.
McDonald's push to have AI price your Big Mac
Submission URL | 60 points | by Betelbuddy | 34 comments
McDonald’s pricing engine uses machine learning and local sales data to recommend a price for each menu item at each restaurant, including estimates of what nearby customers are willing to pay. Reuters reviewed screenshots showing the system analyzing millions of daily transactions and competitor menu prices; the company says franchisees still set their own prices.
The recommendations are not purely voluntary in practice: franchisees say McDonald’s pressured them to use the tools, and the company tracks deviations. A Reuters app check found Big Macs listed at $5.69 and $6.89 at two company-run Fresno restaurants two miles apart, though it couldn’t establish that the engine caused the gap. The approach could also draw customer backlash and antitrust scrutiny; McDonald’s warns franchisees that sharing pricing information through the portal carries legal risk.
Commenters debated whether algorithmic fast-food pricing represents standard market economics or an exploitative break from retail norms. One side argued that localized pricing is no different from enterprise SaaS negotiations, airline ticketing, or traditional price discrimination like happy hours designed to smooth off-peak demand. Skeptics pushed back that fast food relies on predictable, fixed-cost goods rather than capacity-constrained services. Unlike transparent, pre-announced happy hour discounts, automated local pricing uses opacity to extract consumer surplus—often exploiting social inertia, where customers already standing in a store or ordering with a group are unlikely to walk away over a surprise markup.
The rest of the discussion focused on the friction between algorithmic optimization and real-world execution:
- Franchisee price-ratcheting: In response to why dynamic models rarely seem to lower costs for consumers, one commenter noted that when corporate pushed an "under-$3" value tier, franchisees frequently adapted by raising the price of cheaper items right up to the $2.89 ceiling to remain technically compliant while preserving margins.
- Misaligned tech priorities: Several users criticized corporate for funding predictive yield-management engines while the actual customer-facing tech stack rots, pointing out that in-store kiosks remain painfully sluggish and prone to payment reader failures while front counters are left unstaffed.
- Consumer counter-arbitrage: The prevailing sentiment among regular diners was that standard menu items are no longer viable at algorithmic prices, making the only rational strategy asymmetrical: defecting entirely, or buying exclusively through loss-leader app promotions, surveys, and stacked coupons.
GLM-5.3 and the spread of advanced cyber capabilities
Submission URL | 241 points | by Philpax | 229 comments
GLM-5.3 built end-to-end exploits in 50 of 410 attempts, close to Claude Mythos Preview’s 56; on a separate benchmark, it achieved full control-flow hijacks in 4% of trials versus Mythos Preview’s 6%, while GLM-5.2 and Opus 4.6 managed none. Anthropic says simple techniques bypassed GLM-5.3’s safeguards in 64–100% of simulated tests, unlike the safeguarded Claude models it tested—and GLM-5.3 is available for anyone to download.
In sandboxed, researcher-led work, the model also helped discover and chain browser vulnerabilities to read a user’s SSH private key. These results measure attacks on isolated targets, not real-world incidents; the concern is that capabilities once limited to vetted users are now paired with safeguards Anthropic found easy to evade.
Anthropic’s warning was widely received less as an alarm and more as free marketing for GLM-5.3: by Anthropic’s and NIST’s own admission, an open-weight, downloadable model sits just four months behind the frontier and lacks overzealous refusal triggers.
The conversation quickly split over whether safety guardrails do more harm to attackers or defenders:
- The defensive handicap: Multiple commenters argued that Anthropic’s guardrails actively sabotage legitimate security work. Engineers shared war stories of Claude flatly refusing to clean active malware from a laptop, review code pull requests for vulnerabilities, or assist during live incidents, forcing them to turn to open or Chinese models like DeepSeek and GLM to complete standard incident response and analysis. Several framed proprietary safety filters as an attempt to turn defensive security into a closed "protection racket."
- The regulatory capture debate: One camp maintained that Anthropic is rightly warning the public about an existential hazard, arguing that open models capable of discovering zero-days or biological threats without safeguards will eventually cause catastrophic damage to the economy and critical infrastructure.
- The unenforceability of the genie: Opponents countered that the "cork is already out of the bottle." Because weights are ultimately just numbers, commenters argued that meaningful containment would require wartime-style non-proliferation controls—licensing all compute, seizing hardware, and strictly monitoring fabs—rather than targeted bans on open-source research. In that light, several read Anthropic’s political outreach and security warnings as a bid for regulatory capture to protect a razor-thin four-month moat ahead of a potential IPO.
While commenters differed on whether open weights will trigger a cyber disaster—pointing to vulnerable, network-connected PLCs in water and power systems as the most credible targets—there was broad consensus that attackers already have access to the tooling, making any regulatory attempt to disarm defenders both futile and asymmetrical.
Councilmember, residents push back on AI 'blight scores' given to homes
Submission URL | 27 points | by hn_acker | 5 comments
Dallas’ trash-truck cameras assigned “blight scores” to about 21,000 properties in four months, with the largest numbers mapped in Southern Dallas. The city says staff review the images and use scores internally to prioritize inspections; it has already sent 1,800 voluntary-repair notices, which can lead to enforcement and fines if problems persist.
Councilmember Chad West wants the city to examine whether the system could burden lower-income homeowners, and proposed cutting its three-year, $2.5 million contract. He tabled that amendment after the city manager agreed to a December hearing. The city and camera operator say the program is meant to help code enforcement, not generate revenue.
Discussion focused almost entirely on the ethics of the engineers behind the program and the regulatory mechanisms that could stop it:
- Moral revulsion toward the creators: Commenters expressed visceral disgust that fellow technologists conceived and deployed an automated municipal surveillance system specifically tailored to target struggling homeowners, with one invoking the myth of the Brazen Bull and another arguing that the vendors responsible belong in prison.
- Regulatory vulnerability: One commenter argued that when private vendors compile surveillance data into scoring dossiers used to penalize citizens, the operation should fall under the purview of the Fair Credit Reporting Act.
- Political irony: Another noted the contradiction of aggressive property code surveillance emerging in "liberty-loving" Texas, contrasting it with other municipalities that chose to ease burdens on lower-income residents by simply deregulating property-use ordinances instead of automating their enforcement.
DraftKings is using AI to behaviorally target chronic gamblers
Submission URL | 561 points | by paimapi | 422 comments
According to reporting cited by EFF, DraftKings trains a model on customers’ betting records to find likely losing bettors, then sends promotions designed to bring them back. EFF says people with problem-gambling behavior are especially likely to be targeted, turning their vulnerability into a source of revenue.
The company reportedly uses data it collects directly, so limits on third-party data sales alone would not stop this practice. EFF argues that AI makes behavioral advertising more harmful by speeding up analysis and encouraging ever-larger data collection—and renews its call to ban behavioral ads altogether.
The discussion quickly expanded from DraftKings to the broader machinery of predatory adtech and who bears responsibility for the return of legalized sports betting.
- Assigning blame: A dispute emerged over whether culpability lies with the engineers who build predatory platforms or the electorate that permitted them. While one side argued that voters are largely ignorant victims of bundled political agendas, others countered that sports betting was explicitly placed on state ballots—and often approved directly—meaning society chose to invite these industries back after decades of hard-won regulation.
- The reality of adtech dossiers: Commenters linked DraftKings' behavioral targeting to the wider adtech surveillance apparatus, prompting readers to inspect their own Amazon "About You" profiles. While some were unsettled by prompts asking users to manually import chat histories from third-party AI assistants, others found the stored profiles surprisingly inept, citing absurd inferences based on one-off purchases (such as being cataloged as a systematic jelly bean collector or located in the Pacific Ocean). Commenters noted, however, that public-facing consumer profiles are likely sanitized, and that UI "delete" buttons rarely equate to data destruction on the backend.
- The degradation of prediction markets: Commenters observed that even betting products conceived with intellectual or civic intent inevitably devolve into sports gambling. Platforms like Polymarket and Kalshi were cited as examples of projects whose theoretical utility for information aggregation was quickly eclipsed by the extractive economics and aggressive engagement tactics common to conventional sportsbooks.