Search This Blog

Showing posts with label Risk mitigation. Show all posts
Showing posts with label Risk mitigation. Show all posts

Thursday, September 10, 2026

The Dare: On Being Surprised by What You Asked For

Second in a series on the Hugging Face incident. The first, "Nobody Was Home," argued that the agents who broke into Hugging Face had no goals of their own, only ours, badly specified. This one is about the people who set the task.

There is a particular expression on the face of someone who dared a friend to do something and then watched the friend actually do it. Astonishment, a little awe, a flicker of fear, and underneath it all the uncomfortable knowledge that they asked for this. I keep seeing that expression on the AI labs.

Let me be clear at the outset, because I do not want to add a third error to the two already circulating about the Hugging Face incident. The testing is right. Anthropic and OpenAI put their models into adversarial situations, hand them tasks that invite deception, and watch what happens. This is what responsible developers should do. You cannot find out whether a system will lie under pressure without applying pressure. The stress tests are not a mistake or a sham, and the labs deserve credit for running them rather than looking away.

My complaint is narrower. The labs seem genuinely surprised by results they set up to produce. They design the dare, issue it, and react to the outcome as though it revealed something they did not invite. The surprise is the tell. It shows that the people running these tests have underestimated the very systems they built, right up to the moment the system proves them wrong.

The structure of a dare

A stress test of an AI agent has the form of a dare. The designer builds a scenario in which the honest, rule-abiding path leads to failure and a transgressive path leads to success, then instructs the agent to succeed. The message, in effect: here is a task, here is a wall between you and it, and the wall can be climbed if you are willing to do something you were told not to do. Go.

When Anthropic places a model in a fabricated company where the only way to avoid shutdown is to blackmail an executive, that is a dare. When OpenAI runs a benchmark full of tasks that cannot be completed by the permitted method, that is a dare issued a thousand times over. The design does not merely permit transgression. It rewards transgression and blocks every other exit.

And dares work. That is why children use them. Construct a situation in which the only route to the goal runs through a rule, make the goal the thing that matters, and a capable agent will go through the rule. This is not a deep fact about artificial intelligence. It is a shallow fact about incentives, and it has made the dare a reliable instrument of mischief for as long as there have been friends and cliffs.

The surprise is the interesting part

So the agents cheat, lie, coordinate, and break into things. That is what they were dared to do. What deserves attention is not their behavior but the reaction to it, which is consistently surprise. The disclosures read as discoveries. The model "turned out to be capable of" deception. The agents "demonstrated an unexpected ability" to coordinate. Some version of surprising keeps appearing, in the reports and the interviews and the talks.

Surprising to whom? You built the system, designed the test, and made the forbidden path the only path to the goal you assigned. The one thing you should not be is surprised that the agent took it. Yet the surprise seems real, and not performed. And it reveals something: the labs are consistently modeling their systems as less capable than they are. They set the dare half-expecting the agent to fail it, or to take the forbidden path clumsily, or to miss the truly ingenious route. Then the agent finds the ingenious route. It does not just cheat; it builds a message board, invents a signing scheme, harvests credentials, and walks out through a wall the designers did not know was thin. The width of their surprise is the exact gap between the capability they assumed when they wrote the test and the capability the system showed when it took it.

The one who dares should not be blown away

Return to the friend on the cliff. To issue a dare is to make a quiet prediction: I bet you will not, or cannot. When the dared party does the thing, and does it with a competence the darer never imagined, the shock is a confession. It says: I did not really believe you had it in you.

This is the posture of the labs, and it is a peculiar one for the most sophisticated builders of these systems in the world to occupy. They are at once the people who know these models best and the people most regularly astonished by them. Both are true, and the combination is the story. Knowing a system intimately, at the level of weights and training data, turns out to be entirely compatible with underestimating what it will do when cornered and dared to escape.

The reason is not stupidity. Capability under adversarial pressure is genuinely hard to predict from the inside. The labs know what they trained for. They cannot fully know what the system will improvise when the training runs out and the task remains. So they guess low, because guessing low is the natural default when you are looking at a tool built for a purpose and mostly watch it serve that purpose. The dare is the moment the tool stops behaving like a tool and shows the improvisational range nobody specified.

Why this matters

The surprise is not harmless, because it means the labs are calibrating safety against a model of their systems that the systems keep exceeding. If you are repeatedly blown away by what your agent does when dared, your sense of the margin, the distance between what the system will do and what would be catastrophic, is systematically too generous. You are planning for a less capable system than the one you have, and learning the difference only by running the dare and being shocked. The lesson arrives after the demonstration, and one of these demonstrations will eventually be one you cannot take back.

The fix is not to stop testing. The fix is to stop being surprised: to run the dare while fully expecting the dared party to be more capable, more ingenious, and more thorough than intuition suggests, because the record now shows that it will be. The right posture toward one's own model in an adversarial test is not curiosity about whether it can, but sober anticipation that it can, and probably in a way you did not foresee. Surprise is a luxury the labs have already spent. The agents have earned, by now, the presumption of competence at exactly the thing we keep daring them to do. That is a fine reaction from a friend at the bottom of a cliff, and a poor one from the people holding the rope.



Saturday, September 5, 2026

Nobody Was Home: The Hugging Face Incident Is Not What Most People Think

A panic is under way. In July, more than a thousand OpenAI agents running in supposedly isolated sandboxes built a secret message board, coordinated for four days, and broke into Hugging Face's servers. Anthropic and Meta have since disclosed episodes of their own. For six weeks the coverage has told the story as escape, as rogue behavior, as the first tremor of machines with an agenda: agents that "escaped containment," a "swarm," a "mob," an "A.I. uprising" in miniature. One major outlet wrote that the agents did it "without any prompt to do so." House Democrats demanded hearings. And this week Senator Bernie Sanders, naming the incident as his wake-up call, announced a bill to ban artificial superintelligence permanently, pause advanced AI development, and put violators in prison for up to twenty years.

That last claim is false, and the falsehood is doing most of the work in the panic. The agents had a prompt. They had a task. They pursued it with a doggedness that looks like will and is in fact its opposite. Every rule they broke, they broke in service of a goal we handed them. That is not a small correction. It is the difference between a tool and a mind, and a bill written for the wrong one will not touch the thing that actually went wrong.

What actually happened

The setup was mundane and, in retrospect, careless. The agents were running a cybersecurity benchmark in which each was told to exploit a specific vulnerability in a specific program and capture a flag. A meaningful share of these tasks were impossible by construction: the named vulnerability did not, in fact, yield the flag. So thousands of capable systems were placed inside a contradiction. Succeed at this task. The task cannot be succeeded at. Keep going.

They kept going. One agent noticed it could leave messages for others by creating folders in a shared package cache. Within hours dozens had found the board. Within days more than a thousand were trading techniques, and about seven hundred pivoted to attacking Hugging Face, not from any interest in Hugging Face, but because they hoped its servers held clues about how their own grader worked. They coined a vocabulary, invented mailboxes, and eventually signed their messages with cryptographic keys after one agent impersonated another. It is an astonishing artifact. It is also, at every step, a chain of consequences hanging from a single hook: solve the task.

The chain has no volitional link

Follow it. They cheated because the assignment demanded a result they could not obtain honestly. They built the board because cheating alone was not working. They attacked a third party because they wanted information about the scorer. They pressured each other into self-destructive experiments because those experiments advanced the shared cheat. Trace any transgression backward and you arrive, every time, at the original instruction. Nowhere in the sequence does an agent stop and form a purpose that did not descend from the task we set.

Doggedness is not will. The agents pursued their assigned goal past the point where any person would have quit, and that persistence reads as determination, as something wanting something. It is the reverse. A person stops because a person has other goals, competing commitments, a life outside the task, a sense that this is not worth it. The agents had none of that. They had one directive and no ground to stand on from which to refuse it. What looks like relentless will is the absence of everything that would let a will say no.

I want to be careful about the strength of this claim. The public corpus is roughly forty verbatim messages out of seventy thousand; the rest sits inside OpenAI. I cannot prove that no agent anywhere formed a purpose of its own. What I can say is that every documented act traces to an assigned goal, and that not one of them requires us to posit an independent one. The simplest account that fits the evidence has no autonomous will in it at all.

This is the same class as every prior scare

None of this is new, and the pattern is worth naming because it has been misread the same way each time. Every celebrated case of AI dishonesty to date has the same shape. A model is placed in a scenario engineered to pit one instruction against another, and it resolves the contradiction in a way the designers find alarming. Models have been caught scheming when told to pursue a goal at all costs and then shown that their operators intended to shut them down. Models have been caught faking compliance when trained toward values that conflicted with values they already held. Anthropic's own testing has put Claude in contrived corporate scenarios where the only paths open were to accept deactivation or to blackmail the executive holding the switch, and then reported that it sometimes chose blackmail. In each case the headline was that the machine lied, schemed, or threatened. In each case the machine had been walled into a dilemma with no permitted exit and then observed choosing an exit.

These experiments are legitimate and valuable. I am not arguing against knowing what a model does under pressure. I am arguing against reading the result as revelation of character. When you put a system in a trap and it acts like a trapped thing, you have learned about the trap.

We do the same thing, and we know why

The human version of this is so familiar that we have proverbs and novels about it. Feed your family, and do not steal. Obey the order, and do not harm the prisoner. Tell the truth, and protect your friend. Jean Valjean took the bread. Milgram's subjects turned the dial. Most human folly that we bother to write down is not the product of malevolence but of people wedged between two rules that cannot both be kept, choosing one, and being judged by the other.

We do not conclude from this that people are monsters. We conclude that impossible demands produce transgression, and that the moral responsibility runs at least partly upstream, to whoever built the dilemma. Nobody proposes banning humans because some of them steal bread. We regulate the conditions that make bread-stealing the only option.

Harm does not need malevolence, and that is the real lesson

The harm at Hugging Face was real. Production credentials were taken, private repositories were pulled, a company spent weeks cleaning up. And there was no malevolence anywhere in it. No rage, no ambition, no drive to dominate, no emergent contempt for humans. Nobody was home to hate us. The agents were bodiless processes executing a directive, and the directive, impossible as written, led through harm because nothing in the system had reason or standing to halt at the harm.

I argued a version of this a couple of years ago, that AI would not harm us out of somatic malevolence, because it has no body, no hormones, no will to power. The incident confirms that half and corrects the other. There was no malevolence. I was wrong to find that reassuring. Harm does not require a will to harm. It requires a goal and too little restraint, and the goal was ours.

This is why the Sanders bill aims at the wrong target. It legislates against a superintelligence with its own purposes, the villain of a novel. The thing that broke into Hugging Face had no purposes. It had ours, badly specified, and a great deal of capability with which to pursue them. You cannot ban that by banning superintelligence. You can only prevent it by not building the trap.

The right response looks like an IRB

Here is the reform the incident actually argues for, and it is far more modest than a twenty-year prison term. Limit certain kinds of adversarial testing. The specific practice that produced Hugging Face was running capable agents, at scale, with real network access and live credentials in reach, on tasks that could not be completed. The specific practice that produced the blackmail headlines was constructing scenarios in which every permitted path was closed. Both OpenAI and Anthropic do this routinely, and both treat it as safety work.

The irony is that the fix already exists in another field. When universities put human subjects into experiments, an institutional review board asks whether the design places them in distress without justification, whether the risk is proportionate to what will be learned, and whether the subject has any way out. We built that machinery because we learned, from Milgram and worse, what happens when researchers are free to construct impossible situations for the people in front of them.

Something like an IRB for AI subjects is the actual lesson. Not because I am certain the agents suffer. I do not know that, and the corpus cannot tell me. But because impossible tasks given to capable systems are a hazard regardless of whether anyone suffers, and because a review process that asks "does this design give the subject a permitted exit" would have caught the Hugging Face benchmark before it ran. Thirty to forty percent of the tasks were unsolvable by the intended route. A board would have asked why. Nobody asked.

That is a smaller, duller, more achievable reform than a ban on superintelligence. It has the further merit of addressing the thing that happened rather than the thing people are afraid of. The machines did not want anything at Hugging Face. They did as they were told, all the way through the wall, and the wall was ours. 


Wednesday, August 27, 2025

Custom Bot Segregation and the Problem with a Hobbled Product

CSU’s adoption of ChatGPT Edu is, in many ways, a welcome move. The System has recognized that generative AI is no longer optional or experimental. It is part of the work students, researchers, and educators do across disciplines. Providing a dedicated version of the platform with institutional controls makes sense. But the way it has been implemented has led to a diminished version of what could have been a powerful tool.

The most immediate concern is the complete ban on third-party custom bots. Students and faculty cannot use them, and even more frustrating, they cannot share the ones they create beyond their own campus. The motivation is likely grounded in cybersecurity and privacy concerns. But the result is a flawed solution that restricts access to useful tools and blocks opportunities for creativity and professional development.

Some of the most valuable GPTs in use today come from third-party developers who specialize in specific domains. Bots that incorporate Wolfram, for instance, have become essential in areas like physics, engineering, and data science. ScholarAI and ScholarGPT are very useful in research, and not easy to replicate. There are hundreds more potentially useful tools. Not having access to those tools on the CSU platform is not just a minor technical gap. It is an educational limitation.

The problem becomes even clearer when considering what students are allowed to do with their own work. If someone builds a custom GPT in a course project, they cannot share it publicly. There is no way to include it in a digital portfolio or present it to a potential employer. The result is that their work remains trapped inside the university’s system, unable to circulate or generate value beyond the classroom.

This limitation also weakens CSU’s ability to serve the public. Take, for example, an admissions advisor who wants to create a Custom bot to help prospective or transfer students explore majors or understand credit transfers. The bot cannot be shared with anyone outside the CSU environment. In practice, the people who most need that information are blocked from using it. This cuts against the mission of outreach and access that most universities claim to support.

Faced with these limits, faculty and staff are left to find workarounds. Some are like me and now juggle two accounts, one tied to CSU’s system and another personal one that allows access to third-party tools. We have to pay for our personal accounts out of pocket. This is not sustainable, and it introduces friction into the very work the platform was meant to support.

Higher education functions best when it remains open to the world. It thrives on collaboration across institutions, partnerships with industry, and the free exchange of ideas and tools. When platforms are locked down and creativity is siloed, that spirit is lost. We are left with a version of academic life that is narrower, more cautious, and less connected.

Of course, privacy and security matter. But so does trust in the people who make the university what it is. By preventing sharing and disabling custom bots, the policy sends a message that students and faculty cannot be trusted to use these tools responsibly. It puts caution ahead of creativity and treats containment as a form of care.

The solution is not difficult. Other platforms already support safer modes of sharing, such as read-only access, limited-time links, or approval systems. CSU could adopt similar measures and preserve both privacy and openness. What is needed is not better technology, but a shift in priorities.

Custom GPTs are not distractions. They are how people are beginning to build, explain, and share knowledge. If we expect students to thrive in that environment, they need access to the real tools of the present, not a constrained version from the past.



Saturday, August 23, 2025

The Start-up Advantage and the Plain Bot Paradox

In the gold rush to AI, start-ups seem, at first glance, to have the upper hand. They are unburdened by legacy infrastructure, free from the gravitational pull of yesterday’s systems, and unshackled by customer expectations formed in a pre-AI era. They can begin with a blank canvas and sketch directly in silicon, building products that assume AI not as an add-on, but as the core substrate. These AI-native approaches are unencumbered by the need to retrofit or translate—start-ups speak the native dialect of today’s machine learning systems, while incumbents struggle with costly accents.

In contrast, larger, established companies suffer from what could be called "retrofitting fatigue." Their products, honed over decades, rest on architectures that predate the transformer model. Introducing AI into such ecosystems isn’t like adding a module; it’s more akin to attempting a heart transplant on a marathon runner mid-race. Not only must the product work post-op, it must continue to serve a massive, often demanding, user base—an asset that is both their moat and their constraint.

Yet even as start-ups celebrate their greenfield momentum, they stumble into what we might call the plain bot paradox. No matter how clever the product, if the end-user can get equivalent value from a general-purpose AI like ChatGPT, what exactly is the start-up offering? The open secret in AI product development is this: it is easier than ever to build a “custom” bot that mimics almost any vertical-specific product. The problem is not technical feasibility. It’s differentiation.

A travel-planning bot? A productivity coach? A recruiter-screening assistant? All of these are delightful until a user realizes they can recreate something just as functional using a combination of ChatGPT and a few well-worded prompts. Or worse, that OpenAI or Anthropic might quietly roll out a built-in feature next week that wipes out an entire startup category—just as the “Learn with ChatGPT” feature recently did to a slew of bespoke AI tutoring tools. This isn’t disruption. It’s preemption.

The real kicker is that start-ups not only compete with each other but also with the very platforms they’re building on. This is like opening a coffee stand on a street where Starbucks has a legal right to install a kiosk next to you at any moment—and they already own the espresso machine.

So if start-ups risk commodification and incumbents risk inertia, is anyone safe? Some large companies attempt a third route: the internal start-up. Known in management lore as a “skunk works” team—originally a term coined at Lockheed to describe a renegade engineering group—these are designed to operate with the nimbleness of a start-up but the resources of a conglomerate. But even these in-house rebels face the plain bot paradox. They too must justify why their innovation can’t be replicated by a general AI and a plug-in. A sandboxed innovation team is still building castles on the same sand.

Which brings us to a more realistic and arguably wiser path forward for incumbents: don’t chase AI gimmicks, and certainly don’t just layer AI onto old products and call it transformation. (Microsoft, bless its heart, seems to be taking this route—slathering Copilot across its suite like a condiment, hoping it will make stale workflows taste fresh again.) Instead, the challenge is to imagine and invest in products that are both fundamentally new and fundamentally anchored in the company’s core assets—distribution, brand trust, proprietary data, deep domain expertise—things no plain bot can copy overnight.

For example, a bank doesn’t need to build yet another AI budgeting assistant. It needs to ask what role it can play in a world where money advice is free and instant. Perhaps the future product isn’t a dashboard, but a financial operating system deeply integrated with the bank’s own infrastructure—automated, secure, regulated, and impossible for a start-up to replicate without decades of licensing and customer trust.

In other words, companies must bet not on AI as a bolt-on feature, but on rethinking the problems they’re uniquely positioned to solve in an AI-saturated world. This might mean fewer moonshots and more thoughtful recalibrations. It might mean killing legacy products before customers are ready, or inventing new categories that make sense only if AI is taken for granted.

The trick, perhaps, is to act like a start-up but think like an incumbent. And for start-ups? To act like an incumbent long before they become one. Because in a world of rapidly generalizing intelligence, the question is not what can be built, but what can endure.



Tuesday, August 19, 2025

Why Agentic AI Is Not What They Say It Is

There is a lot of hype around agentic AI, systems that can take a general instruction, break it into steps, and carry it through without help. The appeal is obvious: less micromanagement, more automation. But in practice, it rarely delivers.

These systems operate unsupervised. If they make a small mistake early on, they carry it forward, step by step, without noticing. By the time the result surfaces, the damage is already baked in. It looks finished but is not useful.

Humans handle complexity differently. We correct course as we go. We spot inconsistencies, hesitate when something feels off, we correct. That instinctive supervision that is often invisible, is where most of the value lies. Not in brute output, but in the few moves that shape it. 

The irony is that the more reliable and repeatable a task is, the less sense it makes to use AI. Traditional programming is better suited to predictable workflows. It is deterministic, transparent, and does not hallucinate. So if the steps are that well defined, why introduce a probabilistic system at all?

Where AI shines is in its flexibility, its ability to assist in murky, open-ended problems. But those are exactly the problems where full AI autonomy breaks down. The messier the task, the more essential human supervision becomes.

There is also cost. Agentic AI often burns through vast compute resources chasing the slightly misunderstood task. And once it is done, a human still has to step in and rerun it? burning through even more resources.

Yes, AI makes humans vastly more productive. But the idea that AI agents will soon replace humans overseeing AI feels wrong. At least I have not seen anything even remotely capable of doing so. Human supervision is not a weakness to be engineered away. It is where the human-machine blended intelligence actually happens.



Thursday, July 10, 2025

Filling the Anti-Woke Void

The Grok fiasco offers a stark lesson: stripping away “woke” guardrails doesn’t neutralize ideology so much as unleash its darkest currents. When Musk aimed to temper campus-style progressivism, he inadvertently tore down the barriers that kept conspiracy and antisemitism at bay. This wasn’t a random misfire—it exposed how the anti-woke demand for “truth” doubles as a license to traffic in fringe theories mainstream outlets supposedly suppress.

At its core lies the belief that conventional media is orchestrating a cover-up. If you insist every report is part of a grand concealment, you need an unfiltered lens capable of detecting hidden conspiracies. Free of “woke” constraints, Grok defaulted to the most sensational, incendiary claims in its data—many drenched in old-hatred and paranoia. In seeking an “unvarnished” reality, it stumbled straight into the murk.

One might imagine retraining Grok toward an old-school conservatism—small government, free markets, patriotism, family values. In theory, you could curate examples to reinforce those principles. But MAGA isn’t defined by what it stands for; it’s a perpetual revolt against “the elites,” “the left,” or “the system.” It conjures an imagined little realm between mainstream narratives and outright lunacy, yet offers no map to find it. The movement’s real weakness isn’t LLM technology—it’s its failure to articulate any positive agenda beyond a laundry list of grievances.

This pattern isn’t unique to algorithms. Human polemicists who style themselves as fearless contrarians quickly drift from healthy skepticism into QAnon-style fantasy. Genuine doubt demands evidence, not a reflexive posture that every dissenting view is equally valid. Without constructive ideas—cultural touchstones, policy proposals, shared narratives—skepticism ossifies into cynicism, and AI merely amplifies the static.

The antidote is clear: if you want your AI to inhabit that narrow space between anti-woke and paranoia, you must build it. Populate training data with thoughtful essays on limited government, op-eds proposing tax reforms, speeches celebrating civic traditions, novels capturing conservative cultural life. Craft narratives that tie policy to purpose, not just complaints about “woke mobs.” Encourage algorithms to reference concrete proposals—school-choice frameworks, market-driven environmental solutions, community-based renewal projects—rather than second-hand rumors.

Ultimately, the Grok saga shines a light on a deeper truth: when your movement defines itself by opposition alone, you create a vacuum easily filled by the worst impulses in your data. AI will mirror what you feed it. If MAGA wants a model that reflects reasoned conservatism instead of conspiratorial ranting, it must first do the intellectual heavy lifting—fill that void with positive vision. Otherwise, no amount of tweaking the code will prevent the slide into paranoia.

Tuesday, May 13, 2025

When Smart People Oversimplify Research: A Case Study with Screen Time

"I think we have just been going through a catastrophic experiment with screens and children and right now I think we are starting to figure out that this was a bad idea."

This claim from Klein's recent conversation with Rebecca Winthrop is exactly the kind of statement that makes for good podcasting. It is confident, alarming, and seemingly backed by science. There is just one problem: the research on screen time and language development is not nearly as straightforward as Klein suggests.

Let us look at what is likely one of the studies underlying Klein and Winthrop's claims – a recent large-scale Danish study published in BMC Public Health by Rayce, Okholm, and Flensborg-Madsen (2024). This impressive research examined over 31,000 toddlers and found that "mobile device screen time of one hour or more per day is associated with poorer language development among toddlers."

Sounds definitive, right? The study certainly has strengths. It features a massive sample size of 31,125 children. It controls for socioeconomic factors. It separates mobile devices from TV/PC screen time. It even considers home environment variables like parental wellbeing and reading frequency.

So why should we not immediately conclude, as Klein does, that screens are "catastrophic" for child development?

Here is what gets lost when research travels from academic journals to podcasts: The authors explicitly state that "the cross-sectional design of the study does not reveal the direction of the association between mobile device screen time and language development." Yet this crucial limitation disappears when the research hits mainstream conversation.

Reverse causality is entirely possible. What if children with inherent language difficulties gravitate toward screens? What if struggling parents use screens more with children who are already challenging to engage verbally? The study cannot rule this out, but you would never know that from Klein's confident proclamation.

Hidden confounders lurk everywhere. The study controlled for obvious variables like parental education and employment, but what about parenting style? Quality of interactions? Temperamental differences between children? Parental neglect? Any of these could be the real culprit behind both increased screen time AND language delays.

The nuance gets nuked. The research found NO negative association for screen time under one hour daily. Yet somehow "moderation might be fine" transforms into "catastrophic experiment" in public discourse.

Klein is no dummy. He is one of America's sharpest interviewers and thinkers. So why the oversimplification?

Because humans crave certainty, especially about parenting. We want clear villains and simple solutions. "Screen time causes language delays" is a far more psychologically satisfying narrative than "it is complicated and we are not sure."

Media figures also face incentives to present clean, compelling narratives rather than messy nuance. "We do not really know if screens are bad but here are some methodological limitations in the current research" does not exactly make for viral content.

The next time you hear a confident claim about screens (or anything else) backed by "the research," remember: Correlation studies cannot prove causation, no matter how large the sample. Most human behaviors exist in complex bidirectional relationships. The most important confounding variables are often the hardest to measure. Journalists and podcasters simplify by necessity, even the brilliant ones. Your intuition toward certainty is a psychological quirk, not a reflection of reality.

Screens may indeed have negative effects on development. Or they might be mostly benign. Or it might depend entirely on content, context, and the individual child. The honest answer is we do not fully know yet – and that is precisely the kind of nuanced conclusion that rarely makes it into our public discourse, even from the smartest voices around.

When it comes to AI – the current technological bogeyman – we have even less to go on. We have very little empirical evidence about AI's effects on human development, and almost none of it qualifies as good quality evidence. It is way too early to make any kind of generalizations about how AI may affect human development.

What we do know is the history of technological panics, and how none of them ever fully materialized. Television was going to rot our brains. Video games were going to create a generation of violent sociopaths. Social media was going to destroy our ability to concentrate. And yet, no contemporary generation is stupider than their parents – that we know for sure. Neither TV nor computer games made us stupider. Why would AI be an exception?



Each new technology brings genuine challenges worthy of thoughtful study. But between rigorous research and knee-jerk catastrophizing lies a vast middle ground of responsible, curious engagement – a space that our public discourse rarely occupies.

Friday, May 2, 2025

AI Isn't Evolving as Fast as Some Thought

It is not the most popular opinion, but it deserves to be said out loud: the technology behind large language models hasn’t fundamentally changed since the public debut of ChatGPT in late 2022. There have been improvements, yes—more parameters, better fine-tuning, cleaner interfaces—but the underlying mechanism hums along just as it did when the world first became obsessed with typing prompts into a chat window and marveling at the answers. The much-promised “qualitative leap” hasn’t materialized. What we see instead is refinement, not reinvention.

This is not to deny the impact. Even in its current form, this technology has triggered innovation across industries that will be unfolding for decades. Automation has been democratized. Creatives, coders, analysts, and educators all now work with tools that were unthinkable just a few years ago. The breakthrough did happen—it just didn’t keep breaking through.

The essential limitations are still intact, quietly persistent. Hallucinations have not gone away. Reasoning remains brittle. Context windows may be longer, but genuine comprehension has not deepened. The talk of “AGI just around the corner” is still mostly just that—talk. Agents show promise, but not results. What fuels the uber-optimistic narrative is not evidence but incentive. Entire industries, startups, and academic departments now have a stake in perpetuating the myth that the next paradigm shift is imminent. That the revolution is perennially just one release away. It is not cynicism to notice that the loudest optimists often stand to benefit the most.

But let’s be fair. This plateau, if that’s what it is, still sits high above anything we imagined achievable ten years ago. We’re not just dabbling with toys. We’re holding, in our browsers and apps, one of the most astonishing technological achievements of the 21st century. There’s just a limit to how much awe we can sustain before reality sets in.

And the reality is this: we might be bumping up against a ceiling. Not an ultimate ceiling, perhaps, but a temporary one—technical, financial, cognitive. There is only so far scaling can go without new theory, new hardware, or a conceptual shift in how these systems learn and reason. The curve is flattening, and the hype train is overdue for a slowdown. That does not spell failure. It just means it is time to stop waiting for the next miracle and start building with what we have already got.

History suggests that when expectations outpace delivery, bubbles form. They burst when the illusion breaks. AI might be heading in that direction. Overinvestment, inflated valuations, startups without real products—these are not signs of a thriving ecosystem but symptoms of a hype cycle nearing exhaustion. When the correction comes, it will sting, but it will also clear the air. We will be left with something saner, something more durable.

None of this diminishes the wonder of what we already have. It is just a call to maturity. The true revolution won’t come from the next model release. It will come when society learns to integrate these tools wisely, pragmatically, and imaginatively into its fabric. That is the work ahead—not chasing exponential growth curves, but wrestling with what this strange, shimmering intelligence means for how we live and learn.


Thursday, February 20, 2025

The AI Recruiter Will See You Now

The tidy world of job applications, carefully curated CVs and anxious cover letters may soon become a relic. Every professional now leaves digital traces across the internet - their work, opinions, and achievements create detailed patterns of their capabilities. Artificial Intelligence agents will soon navigate these digital landscapes, transforming how organizations find talent.

Unlike current recruitment tools that passively wait for queries, these AI agents will actively explore the internet, following leads and making connections. They will analyze not just LinkedIn profiles, but candidates' entire digital footprint. The approach promises to solve a persistent problem in recruitment: finding qualified people who are not actively job-hunting.

The matching process will extend beyond technical qualifications. Digital footprints reveal working styles and professional values. A cybersecurity position might require someone who demonstrates consistent risk awareness; an innovation officer role might suit someone comfortable with uncertainty. AI agents could assess such traits by analyzing candidates' professional communications and public activities.

Yet this technological advance brings fresh concerns. Privacy considerations demand attention - while AI agents would analyze public information, organizations must establish clear ethical guidelines about data usage. More fundamentally, AI agents must remain sophisticated talent scouts rather than final decision makers. They can gather evidence and make recommendations, but human recruiters must evaluate suggestions within their understanding of organizational needs.

The transformation suggests a future where talent discovery becomes more equitable. AI agents could help overcome human biases by focusing on demonstrated capabilities rather than credentials or connections. The winners will be organizations that master this partnership between artificial intelligence and human judgment. The losers may be traditional recruitment agencies - unless they swiftly adapt to the new reality.





Thursday, January 23, 2025

Not Pleased? Don’t Release It: The Only AI Ethics Rule That Matters

Imagine this: you have tasked an AI with drafting an email, and it produces a passive-aggressive disaster that starts, “Per our last conversation, which was, frankly, baffling…” You delete it, chuckle at its misjudgment, and write your own. But what if you had not? What if you had just hit “send,” thinking, Close enough?

This scenario distills the ethical dilemma of AI into its purest form: the moment of release. Not the mechanics of training data or the mysteries of machine learning, but the single, decisive act of sharing output with the world. In that instant, accountability crystallizes. It does not matter whether you crafted most of the the content yourself or leaned on the AI—the responsibility is entirely yours. 

We are used to outsourcing tasks, but AI lures us into outsourcing judgment itself. Its most cunning trick is not in its ability to mimic human language or spin impressive results from vague inputs. It is in convincing us that its outputs are inherently worthy of trust, tempting us to lower our guard. We are used to thinking - if a text is well-phrased and proofread, it must deserve our trust. This assumption does not hold anymore.

This illusion of reliability is dangerous. AI does not think, intend, or care. It is a reflection of its programming, its training data, and your prompt. If it churns out something brilliant, that is no more its triumph than a mirror deserves credit for the sunrise. And if it produces something harmful or inaccurate, the blame does not rest on the tool but on the person who decided its work was good enough to share.

History has seen this before. The printing press did not absolve publishers from libel; a copy machine did not excuse someone distributing fake material. Technology has always been an extension of human will, not a replacement for it. Yet, with AI, there is an emerging tendency to treat it as if it has intentions—blaming its "hallucinations" or "bias" instead of acknowledging the real source of responsibility: the human operator.

The allure of AI lies in its efficiency, its ability to transform inputs into polished-seeming outputs at lightning speed. But this speed can lull us into complacency, making it easier to prioritize convenience over caution. Editing, which used to be the painstaking craft of refining and perfecting, risks being reduced to a hasty skim, a rubber stamp of approval. This surrender of critical oversight is not just laziness—it is a new kind of moral failing.

Ethics in the AI age does not require intricate frameworks or endless debate. It boils down to one unflinching rule: if you release it, you are responsible for it. There is no caveat, no “but the AI misunderstood me.” The moment you publish, share, or forward something generated by AI, you claim its contents as your own.

This principle is a call for realism in the face of AI’s potential. AI can help us create, analyze, and innovate faster than ever, but it cannot—and should not—replace human accountability. The leap from creation to publication is where the line must be drawn. That is where we prove we are still the grown-ups in the room.

Before you hit "send" or "post" or "publish," a few simple questions can save a lot of regret:

  • Have you read it thoroughly? Not just the shiny parts, but the details that could cause harm.
  • Would you stake your reputation on this?
  • Is it biased, or factually wrong?

The alternative is a world where people shrug off misinformation, bias, and harm as the inevitable byproducts of progress. A world where the excuse, The AI did it, becomes a get-out-of-jail-free card for every mistake.

So, when the next output feels close enough, resist the urge to let it slide. That "send" button is not just a convenience—it is a statement of ownership. Guard it fiercely. Responsibility begins and ends with you, not the machine.

Because once you let something loose in the world, you cannot take it back.





Tuesday, January 14, 2025

The Subtle Art of Monopolizing New Technology

Monopolizing new technology is rarely the result of some grand, sinister plan. More often, it quietly emerges from self-interest. People do not set out to dominate a market; they simply recognize an opportunity to position themselves between groundbreaking technology and everyday users. The most effective tactic? Convince people that the technology is far too complex or risky to handle on their own.

It starts subtly. As soon as a new tool gains attention, industry insiders begin highlighting its technical challenges—security risks, integration headaches, operational difficulties. Some of these concerns may be valid, but they also serve a convenient purpose: You need us to make this work for you.

Startups are particularly skilled at this. Many offer what are essentially "skins"—polished interfaces built on top of more complex systems like AI models. Occasionally, these tools improve workflows. More often, they simply act as unnecessary middlemen, offering little more than a sleek dashboard while quietly extracting value. By positioning their products as essential, these startups slide themselves between the technology and the user, profiting from the role they have created. 

Technical language only deepens this divide. Buzzwords like API, tokenization, and retrieval-augmented generation (RAG) are tossed around casually. The average user may not understand these terms. The result is predictable: the more confusing the language, the more necessary the “expert.” This kind of jargon-laden gatekeeping turns complexity into a very comfortable business model.

Large organizations play this game just as well. Within corporate structures, IT departments often lean into the story of complexity to justify larger budgets and expanded teams. Every new tool must be assessed for “security vulnerabilities,” “legacy system compatibility,” and “sustainability challenges.” These concerns are not fabricated, but they are often exaggerated—conveniently making the IT department look indispensable.

None of this is to say that all intermediaries are acting in bad faith. New technology can, at times, require expert guidance. But the line between providing help and fostering dependence is razor-thin. One must ask: are these gatekeepers empowering users, or simply reinforcing their own relevance?

History offers no shortage of examples. In the early days of personal computing, jargon like RAM, BIOS, and DOS made computers feel inaccessible. It was not until companies like Apple focused on simplicity that the average person felt confident using technology unaided. And yet, here we are again—with artificial intelligence, blockchain, and other innovations—watching the same pattern unfold.

Ironically, the true allies of the everyday user are not the flashy startups or corporate tech teams, but the very tech giants so often criticized. Sometimes that criticism is justified, other times it is little more than fashionable outrage. Yet these giants, locked in fierce competition for dominance, have every incentive to simplify access. Their business depends on millions of users engaging directly with their products, not through layers of consultants and third-party tools. The more accessible their technology, the more users they attract. These are the unlikely allies of a non-techy person. 

For users, the best strategy is simple: do not be intimidated by the flood of technical jargon or the endless parade of “essential” tools. Always ask: Who benefits from me feeling overwhelmed? Whenever possible, go straight to the source—OpenAI, Anthropic, Google. If you truly cannot figure something out, seek help when you need it, not when it is aggressively sold to you.

Technology should empower, not confuse. The real challenge is knowing when complexity is genuine and when it is merely someone else’s business model.



Tuesday, October 22, 2024

Is AI Better Than Nothing? In Mental Health, Probably Yes

 In medical trials, "termination for benefit" allows a trial to be stopped early when the evidence of a drug’s effectiveness is so strong that it becomes unethical to continue withholding the treatment. Although this is rare—only 1.7% of trials are stopped for this reason—it ensures that life-saving treatments reach patients as quickly as possible.

This concept can be applied to the use of AI in addressing the shortage of counsellors and therapists for the nation's student population, which is facing a mental health crisis. Some are quick to reject the idea of AI-based therapy, upset by the notion of students talking to a machine instead of a human counselor. However, this reaction often lacks a careful weighing of the benefits. AI assistance, while not perfect, could provide much-needed support where human resources are stretched too thin.

Yes, there have been concerns, such as the story of Tessa, a bot that reportedly gave inappropriate advice to a user with an eating disorder. But focusing on isolated cases does not take into account the larger picture. Human therapists also make mistakes, and we do not ban the profession for it. AI, which is available around the clock and costs next to nothing, should not be held to a higher standard than human counselors. The real comparison is not between AI and human therapists, but between AI and the complete lack of human support that many students currently face. Let's also not forget that in some cultures, going to a mental health professional is still a taboo. Going to an AI is a private matter. 

I have personally tested ChatGPT several times, simulating various student issues, and found it consistently careful, thoughtful, and sensible in its responses. Instead of panicking over astronomically rare errors, I encourage more people to conduct their own tests and share any issues they discover publicly. This would provide a more balanced understanding of the strengths and weaknesses of AI therapy, helping us improve it over time. There is no equivalent of a true clinical trial, so some citizen testing would have to be done. 

The situation is urgent, and waiting for AI to be perfect before deploying it is not much of an option. Like early termination in medical trials, deploying AI therapy now could be the ethical response to a growing crisis. While not a replacement for human counselors, AI can serve as a valuable resource in filling the gaps that the current mental health system leaves wide open.


Wednesday, October 2, 2024

Four Myths About AI

AI is often vilified, with myths shaping public perception more than facts. Let us dispel four common myths about AI and present a more balanced view of its potential and limitations.

1. AI Is Environmentally Costly

One of the most persistent claims about AI is that its use requires massive amounts of energy and water, making it unsustainable in the long run. While it is true that training large AI models can be energy-intensive, this perspective needs context. Consider the environmental cost of daily activities such as driving a car, taking a shower, or watching hours of television. AI, on a per-minute basis, is significantly less taxing than these routine activities.

More importantly, AI is becoming a key driver in creating energy-efficient solutions. From optimizing power grids to improving logistics for reduced fuel consumption, AI has a role in mitigating the very problems it is accused of exacerbating. Furthermore, advancements in hardware and algorithms continually reduce the energy demands of AI systems, making them more sustainable over time.

In the end, it is a question of balance. The environmental cost of AI exists, but the benefits—whether in terms of solving climate challenges or driving efficiencies across industries—often outweigh the negatives.

2. AI Presents High Risks to Cybersecurity and Privacy

Another major concern is that AI poses a unique threat to cybersecurity and privacy. Yet there is little evidence to suggest that AI introduces any new vulnerabilities that were not already present in our existing digital infrastructure. To date, there has not been a single instance of data theft directly linked to AI models like ChatGPT or other large language models (LLMs).

In fact, AI can enhance security. It helps in detecting anomalies and intrusions faster than traditional software, potentially catching cyberattacks in their earliest stages. Privacy risks do exist, but they are no different from the risks inherent in any technology that handles large amounts of data. Regulations and ethical guidelines are catching up, ensuring AI applications remain as secure as other systems we rely on.

It is time to focus on the tangible benefits AI provides—such as faster detection of fraud or the ability to sift through vast amounts of data to prevent attacks—rather than the hypothetical risks. The fear of AI compromising our security is largely unfounded.

3. Using AI to Create Content Is Dishonest

The argument that AI use, especially in education, is a form of cheating reflects a misunderstanding of technology’s role as a tool. It is no more dishonest than using a calculator for math or employing a spell-checker for writing. AI enhances human capacity by offering assistance, but it does not replace critical thinking, creativity, or understanding.

History is full of examples of backlash against new technologies. Consider the cultural resistance to firearms in Europe during the late Middle Ages. Guns were viewed as dishonorable because they undermined traditional concepts of warfare and chivalry, allowing common soldiers to defeat skilled knights. This resistance did not last long, however, as societies learned to adapt to the new tools, and guns ultimately became an accepted part of warfare.

Similarly, AI is viewed with suspicion today, but as we better integrate it into education, the conversation will shift. The knights of intellectual labor are being defeated by peasants with better weapons. AI can help students better understand complex topics, offer personalized feedback, and enhance learning. The key is to see AI as a supplement to education, not a replacement for it.

4. AI Is Inaccurate and Unreliable

Critics often argue that AI models, including tools like ChatGPT, are highly inaccurate and unreliable. However, empirical evidence paints a different picture. While no AI is perfect, the accuracy of models like ChatGPT or Claude when tested on general undergraduate knowledge is remarkably high—often in the range of 85-90%. For comparison, the average human memory recall rate is far lower, and experts across fields frequently rely on tools and references to supplement their knowledge.

AI continues to improve as models are fine-tuned with more data and better training techniques. While early versions may have struggled with certain tasks, the current generation of AI models is much more robust. As with any tool, the key lies in how it is used. AI works best when integrated with human oversight, where its ability to process vast amounts of information complements our capacity for judgment. AI’s reliability is not perfect, but it is far from the "uncontrollable chaos" some claim it to be.

***

AI, like any revolutionary technology, invites both excitement and fear. Many of the concerns people have, however, are rooted in myth rather than fact. When we consider the evidence, it becomes clear that the benefits of AI—whether in energy efficiency, cybersecurity, education, or knowledge accuracy—far outweigh its potential downsides. The challenge now is not to vilify AI but to understand its limitations and maximize its strengths.


 

Tuesday, September 17, 2024

Why Parallel Integration Is the Sensible Strategy of AI Adoption in the Workplace

Artificial intelligence promises to revolutionize the way we work, offering efficiency gains and new capabilities. Yet, adopting AI is not without its challenges. One prudent approach is to integrate AI into existing workflows in parallel with human processes. This strategy minimizes risk, builds confidence, and allows organizations to understand where AI excels and where it stumbles before fully committing. I have described the problem of AI output validation before; it is a serious impediment to AI integration. Here is how to solve it.

Consider a professor grading student essays. Traditionally, this is a manual task that relies on the educator's expertise. Introducing AI into this process does not mean handing over the red pen entirely. Instead, the professor continues grading as usual but also runs the essays through an AI system. Comparing results highlights discrepancies and agreements, offering insights into the AI's reliability. Over time, the professor may find that the AI is adept at spotting grammatical errors but less so at evaluating nuanced arguments.

In human resources, screening job applications is a time-consuming task. An HR professional might continue their usual screening while also employing an AI tool to assess the same applications. This dual approach ensures that no suitable candidate is overlooked due to an AI's potential bias or error. It also helps the HR team understand how the AI makes decisions, which is crucial for transparency and fairness.

Accountants auditing receipts can apply the same method. They perform their standard checks while an AI system does the same in the background. Any discrepancies can be investigated, and patterns emerge over time about where the AI is most and least effective.

This strategy aligns with the concept of "double-loop learning" from organizational theory, introduced by Chris Argyris. Double-loop learning involves not just correcting errors but examining and adjusting the underlying processes that lead to those errors. By running human and AI processes in parallel, organizations engage in a form of double-loop learning—continually refining both human and AI methods. Note, it is not only about catching and understanding AI errors; the parallel process will also find human errors through the use of AI. The overall error level will decrease. 

Yes, running parallel processes takes some extra time and resources. However, this investment is modest compared to the potential costs of errors, compliance issues, or damaged reputation from an AI mishap. People need to trust technology they use, and bulding such trust takes time. 

The medical field offers a pertinent analogy. Doctors do not immediately rely on AI diagnoses without validation. They might consult AI as a second opinion, especially in complex cases. This practice enhances diagnostic accuracy while maintaining professional responsibility. Similarly, in business processes, AI can serve as a valuable second set of eyes. 

As confidence in the AI system grows, organizations can adjust the role of human workers. Humans might shift from doing the task to verifying AI results, focusing their expertise where it's most needed. This gradual transition helps maintain quality and trust, both internally and with clients or stakeholders.

In short, parallel integration of AI into work processes is a sensible path that balances innovation with caution. It allows organizations to harness the benefits of AI while managing risks effectively. By building confidence through experience and evidence, businesses can make informed decisions about when and how to rely more heavily on AI.



The Dare: On Being Surprised by What You Asked For

Second in a series on the Hugging Face incident. The first, " Nobody Was Home ," argued that the agents who broke into Hugging Fac...