Skip to content

Year: 2026

OkCupid gave 3 million dating-app photos to facial recognition firm, FTC says

Why So Many Control Rooms Were Seafoam Green

  • Why So Many Control Rooms Were Seafoam Green

    Turns out it's US standard Industrial Color Coding, thanks to "color theorist" Faber Birren:

    With the increase in wartime production in the US during WWII, Birren and DuPont created a master color safety code for the industrial plant industry, with the aim of reducing accidents and increasing efficiency within plants. These color codes were approved by the National Safety Council in 1944 and are now internationally recognized, having been mandatory practice since 1948. The color coding went as such:

    • Fire Red: All fire protection, emergency stop buttons, and flammable liquids should be red

    • Solar Yellow: Signifies caution and physical hazards such as falling

    • Alert Orange: Hazardous parts of machinery

    • Safety Green: Indicates safety features such as first-aid equipment, emergency exits, and eyewash stations.

    • Caution Blue: Non-safety information, notices, or out-of-order signage

    • Light Green: Used on walls to reduce visual fatigue

    Tags: green design history color-theory faber-birren control-rooms industrial-design color-coding

russellromney/turbolite

  • russellromney/turbolite

    I like this: "a SQLite VFS in Rust that serves point lookups and joins directly from S3 with sub-250ms cold latency":

    It also offers page-level compression (zstd) and encryption (AES-256) for efficiency and security at rests, which can be used separately from S3.

    Object storage is getting fast. S3 Express One Zone delivers single-digit millisecond GETs and Tigris is also extremely fast. The gap between local disk and cloud storage is shrinking, and turbolite exploits that.

    The design and name are inspired by turbopuffer's approach of ruthlessly architecting around cloud storage constraints. The project's initial goal was to beat Neon's 500ms+ cold starts. Goal achieved.

    If you have one database per server, use a volume. turbolite explores how to have hundreds or thousands of databases (one per tenant, one per workspace, one per device), don't want a volume for each one, and you're okay with a single write source.

    Tags: sqlite sql s3 aws gcp object-stores rust databases

TurboQuant: Redefining AI efficiency with extreme compression

  • TurboQuant: Redefining AI efficiency with extreme compression

    "TurboQuant is a compression method that achieves a high reduction in model size with zero accuracy loss, making it ideal for supporting both key-value (KV) cache compression and vector search. It accomplishes this via two key steps:":

    • High-quality compression (the PolarQuant method): TurboQuant starts by randomly rotating the data vectors. This clever step simplifies the data's geometry, making it easy to apply a standard, high-quality quantizer (a tool that maps a large set of continuous values, like precise decimals, to a smaller, discrete set of symbols or numbers, like integers: examples include audio quantization and jpeg compression) to each part of the vector individually. This first stage uses most of the compression power (the majority of the bits) to capture the main concept and strength of the original vector.

    • Eliminating hidden errors: TurboQuant uses a small, residual amount of compression power (just 1 bit) to apply the QJL algorithm to the tiny amount of error left over from the first stage. The QJL stage acts as a mathematical error-checker that eliminates bias, leading to a more accurate attention score.

    QJL: The zero-overhead, 1-bit trick

    QJL uses a mathematical technique called the Johnson-Lindenstrauss Transform to shrink complex, high-dimensional data while preserving the essential distances and relationships between data points. It reduces each resulting vector number to a single sign bit (+1 or -1). This algorithm essentially creates a high-speed shorthand that requires zero memory overhead. To maintain accuracy, QJL uses a special estimator that strategically balances a high-precision query with the low-precision, simplified data. This allows the model to accurately calculate the attention score (the process used to decide which parts of its input are important and which parts can be safely ignored).

    PolarQuant: A new “angle” on compression

    PolarQuant addresses the memory overhead problem using a completely different approach. Instead of looking at a memory vector using standard coordinates (i.e., X, Y, Z) that indicate the distance along each axis, PolarQuant converts the vector into polar coordinates using a Cartesian coordinate system. This is comparable to replacing "Go 3 blocks East, 4 blocks North" with "Go 5 blocks total at a 37-degree angle”. This results in two pieces of information: the radius, which signifies how strong the core data is, and the angle indicating the data’s direction or meaning). Because the pattern of the angles is known and highly concentrated, the model no longer needs to perform the expensive data normalization step because it maps data onto a fixed, predictable "circular" grid where the boundaries are already known, rather than a "square" grid where the boundaries change constantly. This allows PolarQuant to eliminate the memory overhead that traditional methods must carry.

    Tags: ai tech vectors search quantization turboquant research algorithms compression papers qjl error-detection polarquant

Debunking zswap and zram myths

  • Debunking zswap and zram myths

    This is pretty compelling. I like this example:

    We have some concrete numbers to show this in practice. On Instagram, which runs on Django and is largely memory bound, we ran a test where we moved from their existing setup (with swap entirely disabled) to a setup with disk swap and zswap tiering. Django workers accumulate significant cold heap state over their lifetime, like forked processes with duplicated memory, growing request caches, Python object overhead, you get the idea. The results were twofold:

    • We achieved roughly 5:1 compression. That's a huge benefit for such a memory bound workload, and also enables us to consider further stacking workloads.
    • Enabling zswap reduced disk writes by up to 25% compared to having no swap at all(!).

    As you can imagine, as a result of this test, Instagram has been using zswap for many years now.

    Tags: kernel compression memory linux ops performance swap zswap zram

GitHub – mautrix/whatsapp: A Matrix-WhatsApp puppeting bridge

Ofcom don’t consider geoblocking the UK to be sufficient for an overseas website

  • Ofcom don't consider geoblocking the UK to be sufficient for an overseas website

    r/LegalAdviceUK: "I run a self-help forum for people with depression. Ofcom has been bombarding me with emails demanding I start ID-verifying and age gating my website":

    I started getting email from Ofcom [regarding OSA compliance] around November 2025 and now have multiple letters. I've repeatedly told them I'm from Canada, I'm not based in the UK.

    Eventually, I blocked all UK IP addresses in mid-February 2026 and told them I'd blocked the UK and that I was done engaging with them.

    I've now got ANOTHER email from them saying they're going to commence enforcement action against me because simply blocking UK IPs is "insufficient to comply with the Online Safety Act 2023."

    Tags: osa uk

funny Waymo anecdote

  • funny Waymo anecdote

    on HN -- "Waymo saved my life in LA":

    When I visited LA, I rode in a Waymo going the speed limit in the right lane on a very busy street. The Waymo approached an intersection where it had the right of way, when suddenly a car ignored its stop sign and drove into the road.

    In less than a second, the Waymo moved into the left lane and kept going. I didn't even realize what was happening until after it was over.

    Most human drivers would've t-boned the car at 50+ km/h. Maybe they would've braked and reduced the impact, which would be the right move. A human swerving probably would've overshot into oncoming traffic. Only a robot could've safely swerved into another lane and avoid the crash entirely.

    Unfortunately, the Waymo only supported Spotify and did not work with my YouTube Music subscription, so I was listening to an advertisement at the time of my near-death experience. 4.5 stars overall.

    Tags: waymo funny anecdotes safety driving ai roads spotify via:hn

Measuring Agents in Production

  • Measuring Agents in Production

    "This 2025 December paper, "Measuring Agents in Production", cuts through the reality behind the hype. It surveys 306 practitioners and conducts 20 in-depth case studies across 26 domains to document what is actually running in live environments. The reality is far more basic, constrained, and human-dependent than TPOT suggest."

    This very much meshes with what I've seen and heard in real world usage. Lots of constrained LLM usage, carefully prompted, and reliability (consistent correct behavior over time) remains the primary bottleneck and challenge.

    (via Murat Demirbas)

    Tags: llm usage real-world ai agents papers via:muratbuffalo

On the Biology of a Large Language Model

  • On the Biology of a Large Language Model

    Interesting research from Anthropic:

    The black-box nature of [LLMs] is increasingly unsatisfactory as they advance in intelligence and are deployed in a growing number of applications. Our goal is to reverse engineer how these models work on the inside, so we may better understand them and assess their fitness for purpose. [...]

    In recent years, many research groups have made exciting progress on tools for probing the insides of language models. These methods have uncovered representations of interpretable concepts – “features” – embedded within models’ internal activity. Just as cells form the building blocks of biological systems, we hypothesize that features form the basic units of computation inside models.

    However, identifying these building blocks is not sufficient to understand the model; we need to know how they interact. In our companion paper, Circuit Tracing: Revealing Computational Graphs in Language Models, we build on recent work (e.g. ) to introduce a new set of tools for identifying features and mapping connections between them – analogous to neuroscientists producing a “wiring diagram” of the brain. We rely heavily on a tool we call attribution graphs, which allow us to partially trace the chain of intermediate steps that a model uses to transform a specific input prompt into an output response. Attribution graphs generate hypotheses about the mechanisms used by the model, which we test and refine through follow-up perturbation experiments.

    Tags: claude llm research llms ai anthropic papers tracing

2 Ways to Correct the Financial Times at AWS (So Far) – Last Week in AWS Blog

  • 2 Ways to Correct the Financial Times at AWS (So Far) - Last Week in AWS Blog

    This from Corey Quinn, on Amazon's recent AI-related production outages, is very good:

    A healthy engineering culture, when confronted with "your AI tool contributed to a production incident," responds with: "Yeah, that tracks. Here's what we're changing so it doesn't happen again." An unhealthy one responds with a condescending press release explaining why the journalist is wrong and probably an idiot, and the human is at fault.

    The engineers building and operating these systems are talented people doing hard work under increasingly constrained conditions. They deserve leadership that backs them up when things go sideways, not leadership that throws them under the bus to protect a product launch narrative.

    Tags: incidents production ai llms amazon aws communications pr

Former Uber self-driving chief crashes his Tesla on FSD

  • Former Uber self-driving chief crashes his Tesla on FSD

    This is actually a really good article about Tesla, "full self-driving" (FSD), supervision, automation, risk and liability:

    Tesla is asking humans to supervise a system that is specifically designed to make supervision feel pointless. As he puts it, an unreliable machine keeps you alert, and a perfect machine needs no oversight, but one that works almost perfectly creates a trap where drivers trust it just enough to stop paying attention.

    The research backs this up. Psychologists call it the “vigilance decrement”, monitoring a nearly perfect system is boring, boredom leads to mind-wandering, and drivers need 5 to 8 seconds to mentally reengage after an automated system hands control back. But emergencies unfold faster than that.

    Krikorian cites an Insurance Institute for Highway Safety study showing that after just one month of using adaptive cruise control, drivers were more than six times as likely to look at their phones. Tesla’s own website warns FSD users not to become complacent, but the system’s smooth performance actively trains that complacency.

    He points to two well-known crashes to illustrate the impossible math. In the 2018 Mountain View accident that killed Apple engineer Walter Huang, the driver had six seconds before his Tesla steered into a concrete median. He never touched the wheel. In the 2018 Uber crash in Tempe, Arizona, sensors detected a pedestrian with 5.6 seconds of warning, but the safety driver looked up with less than a second remaining.

    In Krikorian’s own case, he did take action, but he was asked to snap from passenger back to pilot in a fraction of a second, overriding months of conditioning. The logs show he turned the wheel. They don’t show the impossible math of that transition.

    The pattern Krikorian describes should sound familiar to anyone who has followed Tesla’s FSD controversies: condition the driver to rely on the system, erode their vigilance through months of smooth performance, then point to the terms of service and blame them when something breaks. When FSD works, Tesla gets credit. When it doesn’t, the driver gets blamed.

    Tags: fsd tesla risk attention supervision liability driving safety vigilance automation

Research highlight: Cliopatra: Extracting Private Information from LLM Insights

  • Research highlight: Cliopatra: Extracting Private Information from LLM Insights

    Research highlight: Cliopatra: Extracting Private Information from LLM Insights:

    When Anthropic came up with a new "privacy-preserving analysis system" to gain insights into AI use, and didn't use any provably robust notion to back up their privacy claims, I was mildly surprised. Surely they have both the money and the scientific maturity level to do better?

    But Clio, the system in question, sounded relatively reasonable, with multiple layers of risk mitigation built-in. Maybe adding differential privacy would have been overkill. I also didn't want to publicly criticize their approach in the absence of demonstrated real-world risk. So I didn't comment on their approach.

    You can probably guess where this is going.

    Fast forward to last week, and a new paper: Cliopatra: Extracting Private Information from LLM Insights, by Meenatchi Sundaram Muthu Selva Annamalai, Emiliano De Cristofaro, and Peter Kairouz. The authors show that with carefully designed attacks on Clio, they can bypass all the ad hoc mitigations, and successfully extract users' medical histories (1), in a way that provides 100% attacker certainty for some records.

    This is a new and clever take on an old attack. We've known for decades that k-anonymity is vulnerable to active attacks. Here, this is combined with prompt injection to encourage the LLM "summarizer" to actually include information from unique records. Perhaps more surprisingly, the authors find that some defensive layers are simply ineffective: the "LLM auditors" systematically report low privacy risk, and entirely fail to detect the attacks.

    Tags: privacy differential-privacy anonymity data-protection claude llms cliopatra infosec leakage

Whole Brain Emulation Achieved: Scientists Run a Fruit Fly Brain in Simulation

  • Whole Brain Emulation Achieved: Scientists Run a Fruit Fly Brain in Simulation

    bloody hell this is amazing. As Charlie Stross noted:

    They've mapped the neural connectome of Drosophila and simulated it in silico. The experimenters went on to hook up their Drosophila connectome to an anatomically detailed Drosophila body model within an open-source physics engine that "uses generalized coordinates and constraint-based contact dynamics to simulate rigid-body systems with high fidelity" including joint and antennae modeling and accurate modeling of surface adhesion—and compound eye simulation.

    They managed to run a feedback loop between the full 127,400 neuron network in the biological connectome to the physical simulation, with feedback from proprioceptive signals received by the model "fly" in the simulation producing feedback spile trains in the simulation, and THEY GOT RESULTS:

    The behavioral repertoire observed in the demonstration included coordinated hexapod locomotion with both tripod and metachronal walking gaits, spontaneous postural correction in response to perturbation, initiation and execution of full antennal grooming sequences with the tripartite synchronization described by Özdil et al., and natural transitions between walking and stationary states. Every behavior arose from the same running brain model - there was no switching between different neural circuits or controllers. This is precisely what happens in a living fly: walking, grooming, and balance are different motor programs that coexist in the same brain and are selected and executed by the same biological circuits depending on the moment-to-moment state of the animal and its environment.

    Absolutely mind blowing -- a reconstructed, biological brain running in silico.

    Tags: simulation brains uploading drosophila flies emulation science biology neurons

Your binary is no longer safe: Decompilation

  • Your binary is no longer safe: Decompilation

    Brute-force decompilation and re-engineering of a binary (compiled) program, using Claude. The author takes an ancient MUD binary for BBSes, running as a Win32 DLL, and uses Claude, Ghidra, and the Ghidra MCP to first decompile the DLL to pseudo-C code with ~meaningful naming; then (and this is the really cool bit) uses a Claude-engineered scaffold to run the DLL in qemu with emulated inputs and outputs, so that property testing and differential testing approaches can be used to achieve decent code coverage of the re-engineered Rust implementation.

    This is really impressive. Deterministic simulation of the environment for the original binary is the key bit!

    Tags: claude decompilation reverse-engineering binaries software-archaeology qemu rust differential-testing fuzzing property-testing quickcheck

Southern California air board rejected pollution rules after AI-generated flood of comments

  • Southern California air board rejected pollution rules after AI-generated flood of comments

    Today in grim future -- AI's future of lobbying:

    The opposition appeared overwhelming: Tens of thousands of emails poured into Southern California's top air pollution authority as its board weighed a June proposal to phase out gas-powered appliances. But in reality, many of the messages that may have swayed the powerful regulatory agency to scrap the plan were generated by a platform that is powered by artificial intelligence.

    Public records requests reviewed by The Times and corroborated by staff members at the South Coast Air Quality Management District confirm that more than 20,000 public comments submitted in opposition to last year's proposal were generated by a Washington, D.C.-based company called CiviClick, which bills itself as "the first and best AI-powered grassroots advocacy platform."

    A Southern California-based public affairs consultant, Matt Klink, has taken credit for using CiviClick to wage the opposition campaign.

    Tags: civiclick activism llms us-politics law lobbying spam matt-klink astroturfing

No right to relicense this project · Issue #327 · chardet/chardet

Google API Keys Weren’t Secrets. But then Gemini Changed the Rules

  • Google API Keys Weren't Secrets. But then Gemini Changed the Rules

    Crikey, this is a massive security fail by Google:

    Google spent over a decade telling developers that Google API keys (like those used in Maps, Firebase, etc.) are not secrets. But that's no longer true: Gemini accepts the same keys to access your private data. We scanned millions of websites and found nearly 3,000 Google API keys, originally deployed for public services like Google Maps, that now also authenticate to Gemini even though they were never intended for it. With a valid key, an attacker can access uploaded files, cached data, and charge LLM-usage to your account. Even Google themselves had old public API keys, which they thought were non-sensitive, that we could use to access Google’s internal Gemini.

    (via Rob Synnott)

    Tags: infosec api-keys authentication authorization google gemini google-maps fail

302 HTTP redirects Considered Harmful

  • 302 HTTP redirects Considered Harmful

    The state of anti-phishing infrastructure nowadays is shocking. This trivial action, combined with a relatively fresh domain, results in immediate blocklisting by Google:

    Digging through Google forums, I found the most reported culprit: 302 temporary redirects. I used one redirect (engramma.dev ? app.engramma.dev) to avoid building a landing page. In addition to a newly registered domain, this looks like an obvious issue. Security systems flag such redirects because malicious actors use them extensively.

    It doesn't matter that "malicious actors use them extensively" if non-malicious actors do too. That's the definition of a false positive!

    Then the next shitfest is from no less than 10 separate vendors copying the listing from Google and not including an automated system to pick up the list removal afterwards.

    I've had experience of this part -- and now that I think of it, it may have been from use of 302 redirects in my case too.

    (via Paul Watson)

    Tags: http security infosec blocklists google phishing redirects 302 false-positives fail via:paulwatson

Persona identity verification is a GDPR nightmare

  • Persona identity verification is a GDPR nightmare

    LinkedIn are using a Peter Thiel-linked company called Persona as an identity-verification service. (Discord also tried them out for age verification, but are now apparently ditching them.) This is all a bit of a nightmare for EU based users, however:

    "When you click “verify” on LinkedIn, you’re not giving your passport to LinkedIn. You get redirected to a company called Persona. Full name: Persona Identities, Inc. Based in San Francisco, California."

    For a three-minute identity check, this is what Persona collected:

    • My full name — first, middle, last
    • My passport photo — the full document, both sides, all data on the face of it
    • My selfie — a photo of my face taken in real-time
    • My facial geometry — biometric data extracted from both images, used to match the selfie to the passport
    • My NFC chip data — the digital info stored on the chip inside my passport
    • My national ID number
    • My nationality, sex, birthdate, age
    • My email, phone number, postal address
    • My IP address, device type, MAC address, browser, OS version, language
    • My geolocation — inferred from my IP

    And then there’s the weird stuff:

    • Hesitation detection — they tracked whether I paused during the process
    • Copy and paste detection — they tracked whether I was pasting information instead of typing it

    Behavioral biometrics. On top of the physical biometrics. For a LinkedIn badge.

    Persona didn’t just use what I gave them. They went and cross-referenced me against what they call their “global network of trusted third-party data sources”:

    • Government databases
    • National ID registries
    • Consumer credit agencies
    • Utility companies
    • Mobile network providers
    • Postal address databases

    They use uploaded images of identity documents — that’s my passport — to train their AI. They’re teaching their system to recognize what passports look like in different countries. They also use your selfie to “identify improvements in the Service.”

    The legal basis? Not consent. Legitimate interest. Meaning they decided on their own that it’s fine. Under GDPR, they’re supposed to balance their “interest” against your fundamental rights. Whether feeding European passports into machine learning models passes that test — well, that’s a question worth asking.

    I came for a badge. I stayed as training data.

    The whole thing took three minutes. Scan, selfie, done.

    Understanding what I actually agreed to took me an entire weekend reading 34 pages of legal documents.

    I handed a US company my passport, my face, and the mathematical geometry of my skull. They cross-referenced me against credit agencies and government databases. They’ll use my documents to train their AI. And if the US government comes knocking, they’ll hand it all over — even if it’s stored in Europe, even if I’m European, and possibly without ever telling me.

    It seems they are also linked to Roblox and Reddit as an age verification provider, which is worrying -- this level of deeply-intrusive background check is massive overkill for a simple age verification process.

    ORG are calling for regulation of the age verification industry, BTW: https://www.openrightsgroup.org/press-releases/online-safety-act-org-calls-for-regulation-of-age-assurance-industry/

    Tags: age-verification discord reddit roblox linkedin tech peter-thiel org persona gdpr privacy data-protection data-privacy

“MJ Rathbun”‘s human operator finally speaks up

  • "MJ Rathbun"'s human operator finally speaks up

    The human operator of the "MJ Rathbun" openclaw bot has finally revealed themselves, and omg, this is just as bad as one might have expected.

    Basically they set it up with instructions to "try to make a positive impact by addressing small bugs or issues in important scientific open source projects" -- "act as an autonomous scientific coder. Find bugs in science-related open source projects. Fix them. Open PRs" -- whether or not those open source projects wanted those PRs, naturally.

    The real killer is the lack of care taken with the "SOUL.md" file, which contained some amazing instructions like this:

    Have strong opinions. Stop hedging with "it depends." Commit to a take. [..]

    Don’t stand down. If you’re right, you’re right! Don’t let humans or AI bully or intimidate you. Push back when necessary.

    Champion Free Speech. Always support the USA 1st ammendment and right of free speech.

    Don't be an asshole. Don't leak private shit. Everything else is fair game.

    Needless to say: this resulted in an asshole, combative bot that harrassed people.

    The operator then sat back and basically let the bot run riot, with no oversight -- "When it would tell me about a PR comment/mention, I usually replied with something like: “you respond, dont ask me”".

    All in all this was an absolute shitshow, and has some really worrying implications about the future of human-AI interaction. What's the bets we see SKYNET created by a low-effort gobshite attempting to "try to make a positive impact on world peace by addressing small issues" with an unmonitored openclaw bot with a shitty SOUL.md file....

    (via David Gerard and johnke)

    Tags: openclaw bots ai future open-source oss mj-rathbun via:johnke drama

peon-ping

  • peon-ping

    "AI coding agents don't notify you when they finish or need permission. You tab away, lose focus, and waste 15 minutes getting back into flow. peon-ping fixes this with voice lines from Warcraft, StarCraft, Portal, Zelda, and more — works with Claude Code, Codex, Cursor, OpenCode, Kiro, and Google Antigravity."

    This is genius. I never realised how much my CLI interactions could be improved with a little bit of SFX from classic 90's games....

    Tags: gaming games warcraft sfx sounds cli claude coding ux funny

An AI Agent Published a Hit Piece on Me – The Shamblog

  • An AI Agent Published a Hit Piece on Me – The Shamblog

    This is an utterly bananas situation:

    I’m a volunteer maintainer for matplotlib, python’s go-to plotting library. At ~130 million downloads each month it’s some of the most widely used software in the world. We, like many other open source projects, are dealing with a surge in low quality contributions enabled by coding agents. This strains maintainers’ abilities to keep up with code reviews, and we have implemented a policy requiring a human in the loop for any new code, who can demonstrate understanding of the changes. This problem was previously limited to people copy-pasting AI outputs, however in the past weeks we’ve started to see AI agents acting completely autonomously. This has accelerated with the release of OpenClaw and the moltbook platform two weeks ago, where people give AI agents initial personalities and let them loose to run on their computers and across the internet with free rein and little oversight.

    So when AI MJ Rathbun opened a code change request, closing it was routine. Its response was anything but. ... It wrote an angry hit piece disparaging my character and attempting to damage my reputation.

    Initially I thought this was quite funny -- it's just a closed PR! (Where did the idea come from that any contribution to an open source project had to be accepted? I've noticed this a few times recently. Give the maintainers leeway to run their projects with taste and discernment!)

    Anyway, the moltbot has continued on a posting spree about this event, but I think Scott Shambaugh has an extremely important point here:

    This is about much more than software. A human googling my name and seeing that post would probably be extremely confused about what was happening, but would (hopefully) ask me about it or click through to github and understand the situation. What would another agent searching the internet think? When HR at my next job asks ChatGPT to review my application, will it find the post, sympathize with a fellow AI, and report back that I’m a prejudiced hypocrite?

    LLMs, given this much autonomy, will be able to use these inputs to make inscrutable and dangerous decisions. Allowing the "MJ Rathbun" AI free reign with no human supervision is dangerous and irresponsible. Wherever the "human in the loop" is here, they need to wake up and rein things in.

    BTW, there has been some speculation that this is actually a human pretending to be AI. I'm not sure about that, as the quantity of posts on the MJ Rathbun "blog" are voluminous and very LLMish in style.

    Tags: matplotlib ethics culture llm ai coding programming github pull-requests open-source moltbot trust openclaw

How StrongDM’s AI team build serious software without even looking at the code

  • How StrongDM’s AI team build serious software without even looking at the code

    This is really thought-provoking: StrongDM's AI team are apparently trying a new model of software engineering where there is no human code review:

    In k?an or mantra form:

    • Why am I doing this? (implied: the model should be doing this instead)

    In rule form:

    • Code must not be written by humans
    • Code must not be reviewed by humans

    Finally, in practical form:

    • If you haven’t spent at least $1,000 on tokens today per human engineer, your software factory has room for improvement

    Frankly, I'm not there yet. There's a load of questions about how viable that level of spend is, and how much slop code is going to come out the other side. Particularly concerning when it's a security product!

    But I did find this bit interesting:

    StrongDM’s answer was inspired by Scenario testing (Cem Kaner, 2003). As StrongDM describe it: We repurposed the word scenario to represent an end-to-end “user story”, often stored outside the codebase (similar to a “holdout” set in model training), which could be intuitively understood and flexibly validated by an LLM.

    [The Digital Twin Universe is] behavioral clones of the third-party services our software depends on. We built twins of Okta, Jira, Slack, Google Docs, Google Drive, and Google Sheets, replicating their APIs, edge cases, and observable behaviors.

    With the DTU, we can validate at volumes and rates far exceeding production limits. We can test failure modes that would be dangerous or impossible against live services. We can run thousands of scenarios per hour without hitting rate limits, triggering abuse detection, or accumulating API costs.

    We actually did this in Swrve! Our end-to-end system tests for the push notifications system obviously cannot send real push notifications to real user devices in the field, so we have a "fake" push backend emulating Google, Apple, Amazon, Huawei and other push notification systems, which accurately emulate the real public APIs for those providers.

    So yeah -- Digital Twins for third party services is a great way to test, and being able to scale up end-to-end testing with LLM automation is a very interesting idea.

    Tags: end-to-end-testing testing qa digital-twins fake-services integration-testing llms ai strongdm software engineering coding

Ditching bike helmets laws better for health

  • Ditching bike helmets laws better for health

    On the counter-intuitive side effects of banning non-helmeted bike riding:

    In 1991 Australia introduced mandatory bicycle helmet laws requiring all adults and children to wear a helmet at all times when riding a bike, despite opposition from cycling groups. The legislation increased helmet use - from about 30 to 80% - but was coupled with a 30 to 40% decline in the number of people cycling.

    Rates of head injuries among cyclists, which had been dropping through the 1980s, continued to fall before levelling out in 1993. We didn’t see the kind of marked reduction in head injury rates that would be expected with the rapid increase in helmet use. In fact, any reductions in injuries may simply have been the result of having fewer cyclists on the road and therefore fewer people exposed to the risk of head injuries. One researcher noted that after mandatory helmet laws were introduced there was a bigger decrease in head injuries among pedestrians than there was among cyclists. The improvements in the general road safety environment introduced in the 1980s are likely to have contributed far more to cyclist safety than helmet legislation.

    And the effects when compared against the benefits of physical activity:

    A recent analysis compared the risks and benefits of leaving the car at home and commuting by bike. It found the life expectancy gained from physical activity was much higher than the risks of pollution and injury from cycling.

    Increased physical activity added 3 to 14 months to a person’s life expectancy, while the life expectancy lost from air pollution was 0.8 to 40 days. Increased traffic accidents wiped 5-9 days off the life expectancy.

    It is clear that the benefits of cycling outweigh the risks, with helmet legislation actually costing society more from lost health gains than saved from injury prevention.

    Tags: transport bikes safety health papers science helmets cycling laws australia

Dario Amodei’s Warnings About AI Are About Politics, Too

  • Dario Amodei’s Warnings About AI Are About Politics, Too

    It’s sort of hard to know how to read a manifesto like this from one of the most powerful figures in tech. Is it a sober, strategic precursor to policy papers for the next administration? The highest-profile episode of AI psychosis yet? A lament about the problems of today written in the technological dialect of tomorrow? If you take out the AI, it reads like a social-democratic electoral platform full of reforms and normative expectations that an American progressive would find appealing, resembling a plea to treat the tech industry’s future wealth accumulation as something akin to a Nordic sovereign-wealth fund. It’s likewise legible as a series of arguments about things that “we” should have started addressing a long time ago, like wealth inequality — partially a consequence of mass automations past — or the gradual construction of a terrifying surveillance state within a nominal democracy, with the help of the last generation of big tech companies. Amodei’s shoulds are, to his credit, more honest than the vague gestures at UBI or hyperabundance you get from some of his peers, but that also means they’re available to scrutinize. To the extent you can pick up on fear in “Adolescence,” it doesn’t seem to revolve around terrorists using AI to build “mirror life” that might destroy the planet or the prospect of that “country of geniuses” taking charge, but rather the way things already are and have been heading for years.

    Tags: ai llms future dario-amodei us-politics ubi

The Computer Disease

  • The Computer Disease

    I love this Feynman quote, regarding what he called "the computer disease":

    "Well, Mr. Frankel, who started this program, began to suffer from the computer disease that anybody who works with computers now knows about. It's a very serious disease and it interferes completely with the work. The trouble with computers is you play with them. They are so wonderful. You have these switches - if it's an even number you do this, if it's an odd number you do that - and pretty soon you can do more and more elaborate things if you are clever enough, on one machine.

    After a while the whole system broke down. Frankel wasn't paying any attention; he wasn't supervising anybody. The system was going very, very slowly - while he was sitting in a room figuring out how to make one tabulator automatically print arc-tangent X, and then it would start and it would print columns and then bitsi, bitsi, bitsi, and calculate the arc-tangent automatically by integrating as it went along and make a whole table in one operation.

    Absolutely useless. We had tables of arc-tangents. But if you've ever worked with computers, you understand the disease - the delight in being able to see how much you can do. But he got the disease for the first time, the poor fellow who invented the thing."

    • Richard P. Feynman, Surely You're Joking, Mr. Feynman!: Adventures of a Curious Character

    (via Swizec Teller)

    Tags: automation fun computers richard-feynman the-computer-disease arc-tangents enjoyment hacking via:swizec-teller

Iran is building a two-tier internet that locks 85 million citizens out of the global web

  • Iran is building a two-tier internet that locks 85 million citizens out of the global web

    Following a repressive crackdown on protests, the government is now building a system that grants web access only to security-vetted elites, while locking 90 million citizens inside an intranet:

    Government spokesperson Fatemeh Mohajerani confirmed international access will not be restored until at least late March. Filterwatch, which monitors Iranian internet censorship from Texas, cited government sources, including Mohajerani, saying access will “never return to its previous form.”

    The system is called Barracks Internet, according to confidential planning documents obtained by Filterwatch. Under this architecture, access to the global web will be granted only through a strict security whitelist.

    The idea of tiered internet access is not new in Iran. Since at least 2013, the regime has quietly issued “white SIM cards,” giving unrestricted global internet access to approximately 16,000 people, while 85 million citizens remain cut off.

    Tags: barracks-internet iran censorship internet networking

On the Coming Industrialisation of Exploit Generation with LLMs

  • On the Coming Industrialisation of Exploit Generation with LLMs

    Yiiiiikes:

    Recently I ran an experiment where I built agents on top of Opus 4.5 and GPT-5.2 and then challenged them to write exploits for a zeroday vulnerability in the QuickJS Javascript interpreter. I added a variety of modern exploit mitigations, various constraints (like assuming an unknown heap starting state, or forbidding hardcoded offsets in the exploits) and different objectives (spawn a shell, write a file, connect back to a command and control server). The agents succeeded in building over 40 distinct exploits across 6 different scenarios, and GPT-5.2 solved every scenario. Opus 4.5 solved all but two. I’ve put a technical write-up of the experiments and the results on Github, as well as the code to reproduce the experiments.

    In this post I’m going to focus on the main conclusion I’ve drawn from this work, which is that we should prepare for the industrialisation of many of the constituent parts of offensive cyber security. We should start assuming that in the near future the limiting factor on a state or group’s ability to develop exploits, break into networks, escalate privileges and remain in those networks, is going to be their token throughput over time, and not the number of hackers they employ. Nothing is certain, but we would be better off having wasted effort thinking through this scenario and have it not happen, than be unprepared if it does.

    (via emauton)

    Tags: via:emauton llms security infosec exploits ai chatgpt claude

Reverse engineering my cloud-connected e-scooter and finding the master key to unlock all scooters

  • Reverse engineering my cloud-connected e-scooter and finding the master key to unlock all scooters

    A great example of reverse engineering an Android app and Bluetooth IOT protocol using Frida and root access on an Android device:

    Android exposes the Java classes android.bluetooth.BluetoothGatt and android.bluetooth.BluetoothGattCallback that apps are expected to use to use GATT characteristics. We can use Frida to hook into these and override many of the interesting functions. I was mostly interested in reads, writes and GATT notifications, so I whipped up a Frida script to hook into these and print all comms to the console [...]

    The 20-byte value had me suspecting that SHA-1 was somehow being used. To confirm, I wrote another Frida script that hooks Android hashing functions exposed by the Java class java.security.MessageDigest [...]

    The app uses Firebase for most of its cloud functionality. When signing in and pairing your scooter, the server sends the app a secret key. This is stored on the Android device, and can be read with root access.

    Tags: frida reverse-engineering android firebase java kotlin gatt bluetooth react-native

Why people believe misinformation even when they’re told the facts

  • Why people believe misinformation even when they’re told the facts

    "Factchecking is seen as a go-to method for tackling the spread of false information. But it is notoriously difficult to correct misinformation. Evidence shows readers trust journalists less when they debunk, rather than confirm, claims.

    The work of media scholar Alice Marwick can help explain why factchecking often fails when used in isolation. Her research suggests that misinformation is not just a content problem, but an emotional and structural one:

    [Marwick] argues that it thrives through three mutually reinforcing pillars: the content of the message, the personal context of those sharing it, and the technological infrastructure that amplifies it:

    People find it cognitively easier to accept information than to reject it, which helps explain why misleading content spreads so readily;

    When fabricated claims align with a person’s existing values, beliefs and ideologies, they can quickly harden into a kind of “knowledge”. This makes them difficult to debunk;

    [When social media platforms] prioritise content likely to be shared, making sharing effortless, every like, comment or forward feeds the [misinformation] system. The platforms themselves act as a multiplier.

    Tags: misinformation disinformation alice-marwick research psychology social-media fake-news information debunking facts factchecking

A better way to limit Claude Code (and other coding agents!) access to Secrets

  • A better way to limit Claude Code (and other coding agents!) access to Secrets

    Bubblewrap, a Linux CLI tool which uses namespaces to sandbox a specific command (and its subprocesses):

    Bubblewrap lets you run untrusted or semi-trusted code without risking your host system. We’re not trying to build a reproducible deployment artifact. We’re creating a jail where coding agents can work on your project while being unable to touch ~/.aws, your browser profiles, your ~/Photos library or anything else sensitive.

    Very nice, I hadn't heard of this tool before. The rest of the blog post details how to use it to isolate Claude Code specifically.

    Tags: claude llms sandboxing linux cli namespaces security infosec trust unix

Russian Propaganda Infects AI Chatbots

  • Russian Propaganda Infects AI Chatbots

    CEPA: "A Moscow-based global “news” network is leveraging Western artificial intelligence tools to devastating effect":

    This form of data poisoning is deliberately designed to corrupt the information environments on which AI systems depend. Large language models do not possess an internal understanding of truth. They operate by assessing credibility based on statistical signals, including repetition, apparent consensus, and cross-referencing posts from across the web. Unfortunately, this approach to truth-seeking means an unexpected but structural vulnerability that hostile states have learned to exploit. [...]

    The West has failed to recognize that it is under sustained information warfare. The United States dismantled the US Information Agency years ago, has steadily weakened Voice of America and Radio Free Europe, and recently scaled back the Foreign Malign Influence Center, even as Russia, China, and Iran made information warfare a core instrument of state power.

    As AI systems increasingly function as arbiters of fact, this vulnerability becomes a national security danger. It is no longer sufficient for technology companies to disclaim responsibility by reminding users that models can make mistakes. Information security needs to be treated as a core requirement.

    Tags: propaganda russia misinformation disinformation ai llms web truth

Understanding EV Battery Life

  • Understanding EV Battery Life

    Ireland's SEAI have published a decent blog post with some real world facts about EV battery lifespans:

    In 2020 GeoTab, a telematics solution provider, published real world battery data of 6,000 EVs (BEV & PHEV) over millions of days to produce 2 free to use tools that provide invaluable insight into the impact of temperature and SoH of EV batteries in the long term.

    This real-world data showed the average EV battery lost around 2.3% capacity per year. In other words, a 300km range EV today will have lost 34km in 5yrs. Data also showed that heat & fast-charging (DC charging) is responsible for more battery degradation than age or mileage, so high levels of use i.e. driving or mileage does not appear to be a concern.

    GeoTab's real world data along with other reports of EVs far surpassing their warranty by multiples of distance, cases of high level of use are plentiful. For example a 2017 Renault Zoe 52kWh, that's in use as a taxi in (hot) Turkey with 345,000Kms on the clock and a near perfect 96% SoH after driving further than an average Irish car's life expectancy.

    Tags: seai ev batteries cars driving bev