1967 stories
·
2 followers

The 50 Percent Problem

1 Share

In 2024, Laurence Holt of the XQ Institute published an essay titled “The 5 Percent Problem,” a term that quickly became common parlance within the ed-tech zeitgeist. Holt argued that when it comes to measuring the student-learning outcomes of various ed-tech programs, the results often only apply to the five percent of students who “used the program as intended. The other 95 percent see minimal gains, if any.”

Holt’s primary focus was on online math programs, and the central concern running throughout his essay involves whether and how students might “get the recommended dosage” of using various ed-tech tools. Whether the problem arises from unmotivated students or unenthused teachers or disparate access to tech or some combination thereof, the challenge when framed this way is getting students to opt in to using a particular technology.

Generative AI does not suffer from this problem. Uniquely perhaps in the history of ed-tech, AI has been broadly embraced by students worldwide, so much so that the phrase “AI is inevitable” has become commonplace in the education discourse. There are pockets of resistance, of course, and over the last six months we’ve seen mounting opposition to AI across multiple vectors, most prominently with data centers. Nonetheless, we all know that students are using AI frequently. “Dosage” is not an issue.

In the nearly four years of time that’s passed since ChatGPT was commercially deployed, however, we have suffered from a lack of high-quality empirical research on the impact of students using generative AI. There are notable exceptions—such as these—but a recent research landscape analysis out of Stanford indicated that, out of more than 800 studies of AI in education, a mere 20 employed true causal measures. Meanwhile, at least one prominent meta-analysis purporting to show massive learning gains stemming from AI has been retracted, though not before being viewed 400,000 times. The field is a mess.

Perhaps this is because studying AI in “the real world” poses serious challenges. We know that AI tools are free and widely available to students. We also know that many are using them after school hours for various purposes, including to help with their schoolwork. But quantifying this is very difficult, because it’s hard to peer into the home lives of children. And it’s equally if not more challenging to connect out-of-school behavior to measurable learning outcomes, the data sets are not easy to link. Accordingly, research on AI’s education impact has tended toward lab studies or small-scale evaluations.

No longer.

Last week, a study titled “The Generative AI Learning Penalty: Evidence from Chinese Secondary Education,” authored by David Strömberg, Victor Lei, and Yanhui Wu, went viral after The Economist published data visualizations from the research paper. The preprint came out in June—not sure how I missed it—and uses data from approximately 27,000 Chinese students in grades seven to 12 to explore a straightforward question that has preoccupied me for several years: “How does self-directed use of generative AI affect cumulative learning over time in ordinary school settings?”

We will get into the details shortly, but here’s the headline summary of the results:

  • Between 2023 and 2025, approximately 80 percent of all students started using generative AI, and approximately 50 percent of all students engaged in “full homework outsourcing,” meaning, they essentially stopped doing their homework independently and used AI instead.

  • Over time, this caused significant learning loss as measured both on monthly closed-book exams and via the comprehensive high-stakes exams that China uses to determine high school and college placement for students. On those tests, by the end of June 2025, overall student performance fell by 24% on the former and 18% on the latter, a massive decline.

  • What’s more, there is no evidence—none—of any corresponding learning benefit arising from students using AI.

  • As such, the researchers state plainly that “our findings show that generative AI, which is likely to become a prevalent technology for education, has a substantial negative impact on student learning.” (My emphasis)

Put simply, we now have rigorous, empirical evidence on the real world, long-term impact of students using AI. This research indicates that AI substantially harms learning for half of all students, what the researchers call the “Generative AI penalty,” and provides no discernible benefit to the other half. Again, by the time this study concluded, 50 percent of students had fully outsourced their homework to generative AI. For these students, when at home, they just stopped thinking about their schoolwork.

I hereby dub this the 50 Percent Problem. Unlike the Five Percent Problem, the 50 Percent Problem is a measure of harm stemming from widespread adoption of AI by students, rather than their failure to use this technology. It is my conservative estimate of how many students are being directly harmed by using these tools.1 As I’ve said repeatedly for several years, generative AI is a tool of cognitive automation. It’s both predictable and tragic that students are learning less because of it.

This new empirical research provides us with a clear measure of the degree of AI’s educational harm.


Let’s turn to the data itself. In what follows, we’ll be looking at the performance of three different groups of students across three different learning activities.

As to the students, we first have the baseline group of student results from the period prior to the introduction of generative AI—these students are labelled “Pre AI,” and charted in green. Next, we have the admirable group of students, all 19 percent of them, who did not use generative AI at any point between 2023 and June 2025 when the study concluded—these are labelled “Never AI,” and charted in blue. Finally, we have the remaining 81 percent of students who adopted AI and used it for their schoolwork—these are labelled “Used AI,” and charted in red. Of note, these students did not all start using AI at the same time, adoption phased in over time (which allowed the researchers to uncover some interesting things that we won’t go into here).2

As to the learning activities, we’ll start by looking at the amount of time that students spent on their homework; the researchers were able to track this because students had to log in online to download their homework and upload it upon completion. Then, we’ll look at how students scored on this same homework. Last but definitely not least, we’ll examine how student performance changed on monthly closed-book exams, as well as China’s very high-stakes high school and college entrance exams.

Both the underlying research paper and The Economist present a series of helpful data visualizations of this, but without the underlying raw numbers. So I contacted the researchers who conducted the study and David Strömberg (the lead author) graciously agreed to provide me with the underlying histogram data I used to create the three animated GIFs you’re about to see.

With that as backdrop, we can explore three specific questions.

1. What is the impact of AI on the amount of time that students spend on their homework?

The results here are unsurprising—AI significantly reduces the time students spend doing homework. On average, both the Pre AI and Never AI groups spent about an hour doing it, across a range of 50 to 80 minutes. In contrast, within the Used AI group, most students spent around 45 minutes, and some even cruised through in 25, presumably the minimum amount of time it takes to cut-and-paste answers out of a chatbot.

This is AI as tool of cognitive automation working exactly as intended.

2. What is the impact of AI on how students score on their homework?

Here again we see that Pre AI students and Never AI students have near-identical results, as we’d expect. But not so with the Used AI students—now, we see a huge shift to the right in purportedly “positive” outcomes, with scores far higher than even the most studious Never AI student managed to achieve. (Note the scores here were normalized to make the average score a 100 (not the maximum), so a score of 130 means “30 percent above the average,” essentially.)

To restate, the Used AI students studied less yet scored better than their peers. That’s a pretty sweet deal for them, but it’s worth reflecting on the broader implications of this within schools. When I talk to students about AI, one thing I hear them say frequently is that they don’t want to be played for chumps (my term, not theirs). Meaning, if they know their classmates are using AI and getting better grades as a result, even those inclined to resist AI may feel trapped into using it, just to keep pace.

It’s a cognitive race to the bottom.

I’ll also add that homework scores are obviously a very imperfect measure of student learning. In my view, a great deal of education research, and certainly the studies often promoted by ed-tech vendors, use comparable “point in time” data of supposed learning akin to what we see here. This is understandable, to a degree—it’s very difficult to information on long-term learning outcomes. But what ultimately matters in education is building durable student knowledge. The seductive danger of AI exposed here is that, by using AI, students may falsely have believed everything was proceeding swimmingly. In this sense, AI fosters a mental masquerade, scores go up as actual learning goes down.

Now to remove the mask.

3. What is the impact of AI on how students perform on closed-book exams?

Here is where the proverbial rubber meets the road.

Consistent with the two previous data sets, we again see near-perfect alignment between the Pre AI and Never AI students on the monthly closed-book exams administered to students. But now the adverse impact of AI is laid bare—just look at that shift to the left. The Used AI students are scoring at levels far lower than their Never AI peers, indeed, many score far lower than anything recorded prior to AI existing.

I’ve labeled this the Deadweight Learning Loss to underscore the volume of harm caused by the generative AI learning penalty. What’s we’re seeing here is unambiguous evidence of a decline in overall student performance that grew over time as students adopted AI. In fact, the average decline was 20 percent with a 1.4 standard deviation (SD). Although it’s far from an apples-to-apples comparison, this vastly exceeds the estimated learning loss in the US after the pandemic (approximately .25 SD in math and .13 SD in reading).

What’s more, over time this Deadweight Learning Loss had significant adverse consequences for students on the comprehensive high-stakes high school and college entrance exams that China administers to determine student placement.3 For students who adopted generative AI two or more years prior to being tested, the estimated negative effects are 24 percent (1.5 SD) for the high school exam, and 18 percent (1.3 SD) for the college exam.

If you are a parent with a child in junior high or high school who uses generative AI, this data should terrify you. This isn’t about academic integrity, it’s about a generation of kids being told AI is “inevitable” and “the future” and acting accordingly, in a societies where we’ve yet to develop firm norms about what’s acceptable to do with these tools within education. The upshot is that students are using AI in ways that are harming their cognitive development and foreclosing their life opportunities. Harms that are not easily remediated, if at all.

This is the 50 percent problem. Generative AI is cognitive cancer and we are doing next to nothing to stop it from spreading.


I’ll now address a few objections.

First, I want to credit The Economist for putting this research on my radar, and for raising broader attention to its harrowing findings. Yet, remarkably, in the brief article accompanying the research data we’ve just covered, the anonymous magazine author suggests that “AI can boost learning productivity but only for those who use the technology intelligently.”4 In similar fashion, Blake Richards, a researcher at Google who works “on the intersection of machine learning and neuroscience,” argued on social media that AI is not really the problem here, because “if you control for how long students spend studying, then students using AI actually perform equal or better” than those that did not.

Motivated reasoning is a powerful force. Here, of course, “controlling” for how long students spend studying erases the major findings of this research. But let’s leave that aside, and probe whether this argument can be justified on its own terms—does this study suggest that AI can boost learning if used “intelligently”?

As best I can tell, this claim is premised on this graph that correlates homework completion time with exam results:

If we ignore all that pesky data on the left side of this chart (which of course we shouldn’t), and just focus on the overlap between the Generative AI students who continued to study for the same duration as the No Generative AI students, it’s true there’s no major gap between them. Of course, note that there’s no discernible benefit to using AI either—apart that is from the spike at the very top end.

So what’s the deal with the spike? Might we cling to it as proof of the “promise” of AI? Well, here’s a good lesson on why one should always be careful eyeballing charts without access to the relevant underlying data, because when I asked Strömberg how many Used AI students persisted in studying for at least 75 minutes, he told me this (via email, with my emphasis):

It is very rare for students who have adopted AI to spend 75 minutes on homework. This occurs for only 20 students, and for each of them only in a single month (0.2% of AI student-month observations). More than five months after AI adoption, we never observe students spending 75 minutes on homework. In fact, in this group, only four students spend more than 65 minutes on homework.

Let’s get real. Given this study involved almost 27,000 students, indexing on this vanishingly small number of studious AI-using students isn’t just putting lipstick on a pig, it’s smearing its body in Revlon from snout to tail. We need to stop pretending there is some massive benefit to learning if only we trained kids properly on how to use AI. It’s fundamentally harmful.

A different and more sophisticated counterargument might proceed along the following lines: This study tracked student usage of AI stemming from its earliest days, when the tools were less capable than they are today. Perhaps relatedly, the researchers here found evidence that it was the early AI adopters most harmed by generative AI, insofar as “the estimated AI learning penalty fell from around 25 percent in early 2023 to around 16 percent by June 2025.” As such, perhaps this study constitutes the “high-water mark” of AI-induced harm, and perhaps the AI learning penalty will continue to drop. Or so we might hope, anyway.

But as the cliche goes, hope is not a strategy, and there are some problems with this counterclaim. For one thing, the rapid evolution of AI cuts both ways—we now have AI companies explicitly marketing “AI agents” to students to complete their tasks for them. Agents are even worse than chatbots from a cognitive development standpoint—at least the latter require an interaction of some sort to produce output. For another, even if the magnitude of the learning penalty continues to diminish, there’s a question of volume, too. Recall that 20 percent of students managed to resist using AI prior to June 2025. Do you think that number has gone up or down since then? I know my bet.

Finally, I can imagine someone saying I’m placing too much weight on this research—it’s just one study from China, after all. On that front, and as a self-sanity check, I asked three PhD education researchers to review the methodology employed, and all three came back with positive reviews (“it’s very good work,” said one). And China is surely the one country in the world that can match the US for AI adoption and enthusiasm, though it may surprise you to learn that AI companies in China disable their products completely during the high-stakes testing periods. (OpenAI, in revealing contrast, heavily promotes ChatGPT on college campuses during finals week in the US.)

Perhaps more importantly, if you are an AI-in-Education Enthusiast, I feel confident in saying that you will not be able to produce research of comparable rigor that shows positive long-term education impact of generative AI in real-world conditions comparable to those here. If there were such evidence, my inbox would be filled with people jamming it down my throat, trust me. And no, this single study from 2024 involving roughly 150 physics students at Harvard is not going to cut it, sorry.


The 50 Percent Problem will not disappear of its own accord. The Edu-Cognoscenti continues to chatter about learning loss related to school closures during the pandemic. Well, the learning loss stemming from generative AI appears more substantial, it’s happening right now, and it may endure for far longer. So to all the policymakers and philanthropists and so-called thought leaders who purport to take education evidence seriously, I ask you: What are you doing to prevent ongoing educational harms of generative AI? Are you doing anything at all?

We are in cognitive crisis.

I’ll close with this. As I was drafting this essay, my friend Dan Willingham published his perspective about when students should use AI. Please read it. Although he’s a tad more measured (or perhaps just realistic) about its role in education, his core contention echoes what I’ve been arguing for several years, and stems from a basic understanding of human cognition:

[T]he point of assignments is the mental processes required to complete them, and the point of the mental processes is learning. That seems to suggest a simple litmus test for the use of AI. Artificial Intelligence tools should not substitute for tasks wherein students would benefit from doing the mental work themselves. Only use AI for what you already know how to do.

A tool that should only be used once you already know something is not a learning tool. We must stop gesturing at AI’s imagined potential, and focus our efforts instead on mitigating the 50 Percent Problem it’s created.

How, you might reasonably ask? It won’t be easy. But I have some emerging ideas, stemming from recent investigations into communities of technological refusal. It’s been quite the learning journey. More soon.


My thanks to David Strömberg for providing his underlying data and to my anonymous academic friends who reviewed the study and this essay—all errors mine and mine alone, of course.

Subscribe now

1

In my view, 50 percent is a conservative estimate because the researchers themselves describe the core problem in more expansive terms: “The negative effects on learning outcomes appear to be mostly driven by the 81 percent of AI-using students, who spend less time on homework than even the fastest non-AI student, receive high homework scores matching the capability of generative AI tools they are using, and yet very low exam scores.” So I was tempted to call this the 81 Percent Problem, but as discussed above, not all of the 81 percent of AI-using students suffered the AI learning penalty (but nor did they gain any meaningful benefit). Ultimately, it’s the half of students using AI to complete their homework that I’m most worried about.

2

Alert, nerdy readers may be wondering why “student observations” is listed on the Y axis rather than just “students.” Answer: because the students groups were not static and evolved in composition as students adopted AI, the researchers used “student-month” as their unit of analysis, meaning, the data for students was parsed by month—e.g., a single individual student might produce 12 separate “observations” over a year within a single subject. I know, it’s wonky.

3

Of course, whether China or any other education system should employ high-stakes testing to this degree is a highly charged topic, but I’m not interested in having that conversation right now, please and thank you.

4

The Economist also cited this recent study involving the use of chatbots with roughly 200 undergraduates at Middlebury College, conducted over two sessions approximately one week apart. This is a perfect example of a point-in-time education study that bears no resemblance to the reality of how most students are actually using AI.

Read the whole story
mrmarchant
14 hours ago
reply
Share this story
Delete

Melon’s Odyssey On The Black Market For Life-Saving Cat Medicine

1 Share

For a month this spring, Melon was truly living up to her name. She had always been a round cat but never quite this gourd-like—small neck billowing out to a majestic corpulence. This sudden plump came with an attitude. She had abandoned her sweet old habits: sitting on the toilet seat to watch me shower, running back and forth across my apartment until she panted, and leaping on my back to scope out unreachable shelves. I chalked this up to the fact that Melon, now 4 years old, was growing up. Maybe our close bond, which had always felt ineffable to me in ways I cannot articulate without sounding like a TikTok cat mystic, was not actually special, but rather the relationship of a child who needed her parent until, one day, she didn't.

There were other warning signs. Melon had become less and less interested in eating. But she'd always been picky, and I had gone through some drastic life changes—a breakup that came with a division of our two cats—to which I assumed she was still adjusting. Her mane looked terrible, limp and stringy and unkempt, but Melon had never been fastidious about her personal grooming. She'd gotten her annual checkup a few months earlier and left with a clean bill of health. Besides, I was busy, distracted by said life changes and preparing for a vacation and a month off work to write a book. I dropped Melon off at my ex's place, splashed around in Puerto Rico, then came home and picked her up. Neither of us had any idea that anything was wrong. Later that week, my friend Elaine, a fellow cat freak, watched as Melon turned her nose up from a fresh can of wet food. "You know," she said, "we learned Hamlet had cancer when he stopped eating."

My mind began to spin. Hamlet was older when he got cancer, I told myself. Maybe her hunger strike was her acting out for me going on vacation. Friday night, Melon wobbled onto the bed with me and narrowed her eyes into slits. I lay down next to her and studied her face. I tossed, turned, and then spirited us to the emergency animal hospital, empty at 1 a.m. They whisked Melon to the back and I read a book, anticipating the relief I would feel when they would send us home soon with an expensive but clean bill of health. But then I finished my book, and started another. By 3 a.m., I had become too worried to read. I paced my waiting-room cell, decorated only with an ominous daguerreotype of a tuxedo cat.



Read the whole story
mrmarchant
18 hours ago
reply
Share this story
Delete

I made a hologram!

1 Share
I made a hologram!

Recently my family gave me a hologram-making kit that they found at an antique mall in Indiana. I was with them when they first spotted it among the alien masks and vintage comics, and when they pointed it out, my first thought was "yeah right there's no way that's real holograms" followed by "oh wow it looks like actual real holograms".

I made a hologram!
Despite how beat-up the box looks, it still had all its original parts and film, still in their original packaging.

As a laser scientist this was basically the perfect gift for me, and they sneaked back to the antique mall to pick it up. Fortunately for me, the antique mall hologram kit was not all that antique, just a few years old, and the film was still good. The company Litiholo in fact still makes these kits.

The kit included a red laser, and a bunch of plastic holders that pieced together to hold the film and the laser in the right positions. A little rectangle the size of a deck of cards indicated where I could put the 3D object that I wanted to turn into a hologram - they supplied some dice so my first hologram would be of something that would probably work. For example, they recommend not trying to make a hologram of plants because they move too much (!) during the five-minute exposures.

I made a hologram!

I've made holograms in the lab with fancy lasers and floating tables and climate-controlled rooms, so I took their instructions seriously when they told me to sit absolutely still and silent during the hologram exposures. The cats were not allowed in the room.

I made a hologram!
This is the setup for recording holograms. The laser is at the upper right and its film is at the lower left.

The key to the hologram setup is that half of the laser beam hits the film directly, while the other half bounces off the dice and then hits the film. When the two halves interfere with each other at the film, they embed the film with information about the shape of the light bouncing off the dice. Then, when I remove the dice, the film recreates what the dice was doing to its half of the laser beam, as if the dice were still there.

The result is a hologram that changes depending on what angle you look at it, just like you were looking around the edges and over the tops of the original dice. For this video I scooted the dice out from behind the film so now there's nothing back there at all.

You can really see that in this video of some cool calcite crystals; the camera swings up behind the film to show that the calcite isn't there anymore.

Obviously I haven't perfected the technique yet (maybe the film is a little old after all), but I'm having fun experimenting with making holograms. What thing (that's about the size of a deck of cards or smaller) would you want to see a hologram made out of?

Read the whole story
mrmarchant
18 hours ago
reply
Share this story
Delete

A Simplified Mental Model of LLMs

1 Share

Introduction

As of now (late 2026) LLM (large language model) technology providers, users, and work products flood the public commons. It therefore makes sense to have even a primitive mechanistic mental model of these technologies. You are forced to have an opinion. Without a mechanism or model one tends to fall into disempowering anthropomorphic language. Some clear thoughts on this can be found here and here.

In this note I would like to try and outline a (very) simplified mental model of LLM mechanics and mechanisms. By “mental model” I mean a cartoon to work through in your mind, not a model of the LLMs as having their own mind. I won’t be teaching the history of LLMs, how to build them, how to use them, or their moral or philosophic implications. I will only try to give a very rough outline how the current (2026) LLMs work.

The mental model I would like to bring you to is:

  • LLMs are implemented as a flow of numeric signals from a limited number of “attention heads” to output text.
  • LLMs are used to realize text transformation and text construction/fabrication. In particular they are approximate plausible “un-censoring” or “un-deletion” functions.

This note will try to make the above two points concrete and clear. After that I will use the model to drive some discussion/speculation.

Here is an example to explain the type of analogy I am hoping to deliver. A common useful mental model for an internal combustion engine car is: it combines air and fuel to produce motive torque, waste heat, and potentially toxic exhaust. This isn’t enough to build a car, but is enough to tell you not to idle one in an enclosed space.

Caveat

I am assuming the current private frontier LLMs are architecturally similar to GPT-4 (2023), but larger and with improvements. So this note is valid for at most such models.

LLM Implementation

Current LLMs take their form from many engineering decisions, three of the most important (in my opinion) being:

  • Large scale use of artificial “neural networks” (also called deep learning architecture or connectionist architecture). I am going to take the Bender/Inie advice and try to use the non-standard, but much less loaded, term “weighted network.”
  • Clever use of a “un-censor the missing word” training procedures (which led to useful tools such as embeddings).
  • Use of bookmark like structures called “attention.” The primitive components of a weighted network do not actually pay “attention” to anything. Instead they imply weights and selective routing of intermediate values. To not lose this distinction, I will use the term “attention heads” instead of “attention.”

We will describe each of these in turn.

Weighted networks

LLMs are implemented in terms of weighted networks. A weighted network is a representation of nested mathematical expressions or formulae, and not a faithful representation of biology. The earliest weighted networks took a number of input signals (say numbers, or voltages) and combined them into one or more output signals. We can imaging the mechanism as in the following diagram.



In this diagram each of the first three volt-meters (v1, v2, v3) represents input signal values, say the length, width, and height of a box. The three knobs (w1, w2, w3; called “weights” or “parameters”) represent how much of each input signal is passed along the arrows to the next node. The complicated vacuum tube amplifier represents a small function of the inputs: in this case 1/(1 + exp(w1 v1 + w2 v2 + w3 v3)). This weighted network converts 3 input values into one output value (itself represented by the last output meter). These nets are not usually realized using electrical components, but as software specifying calculations in GPUs, TPUs, or NPUs.

The knob settings are called the parameters or weights. The knob-settings are picked in a processed called “training” where we adjust the knobs until the weighted network’s outputs are very close to specified results for a great number of training examples. A training example is a pair of an input example (in this case 3 numbers) and the desired output (in this case one number). Even for very restricted topologies, training can be difficult. For the right choice of the weights (knob settings) this circuit may imitate enough of the example input/output pairs to approximate a useful function. However, researchers have been able to implement and train variations of these networks since the 1950s (ref). It is rumored to have cost around $100 million to train GPT-4 (ref). Current LLMs are much larger and more expensive than that.

The choice of the layout of the circuit is called the “topology” of the weighted network. Our small net’s topology has a feature typical to weighted networks: we can sort the nodes into layers and each layer only connects to later layers (never back or laterally). Our example weighted network consists of one node in one layer with 3 knobs or weights. GPT-4 (the state of the art back in 2023) was thought to have millions of input nodes, possibly billions of intermediate nodes, tens of thousands of output nodes, and about 1.8 trillion weights (knobs or parameters) arranged in possibly 120 layers (ref).

Some things to notice is: even a large weighted network is much weaker than a cheap computer.

  • It has no scratch-pad or short-term memory. When we change its inputs, its output changes independent of where the inputs used to be. Patterns from the training data determine the weights (or knob-settings), but once training is over the knobs are not moved again.
  • No outputs are routed back to earlier portions of the network. This means the calculation can not repeat or iterate steps.

One could implement variations that don’t have the above shortcomings (such as recurrent weighted networks, or trying online reinforcement learning ideas). However current LLMs are thought not to depend heavily on such techniques, as they make training much more expensive for little realized benefit. If a weighted network were modeling biology it would have to have features like the above (as biological neurons are not strictly in layers, and do seem to carry mutable state), but the current known engineering trade-offs are against such features so they tend not to be used.

Clever un-censor training

One hot encoding- the un-clever step

LLMs are demonstrated processing text, not values or numbers as our earlier weighted network did. Researchers adapt weighted networks to text with a very brutal idea called “one hot encoding.” Let’s approach this using an example.

Suppose we with to build a weighted network that works over 9 word utterances from a dictionary of 14 words or tokens. One such utterance is “the quick brown fox jumped over the lazy dog.” To build our adapted net we would build a 9 row (one row for each word in our utterance) by 14 column (one column for each work in our dictionary) paddle-switch array that applies +10 volts where a switch is on and 0 volts where off such as the following.

a

brown

dog

dug

earnest

fox

jumped

lazy

over

quick

red

slow

the

under

the a brown dog dug earnest fox jumped lazy over quick red slow the under
quick a brown dog dug earnest fox jumped lazy over quick red slow the under
brown a brown dog dug earnest fox jumped lazy over quick red slow the under
fox a brown dog dug earnest fox jumped lazy over quick red slow the under
jumped a brown dog dug earnest fox jumped lazy over quick red slow the under
over a brown dog dug earnest fox jumped lazy over quick red slow the under
the a brown dog dug earnest fox jumped lazy over quick red slow the under
lazy a brown dog dug earnest fox jumped lazy over quick red slow the under
dog a brown dog dug earnest fox jumped lazy over quick red slow the under

The voltages from these switches feeds a larger weighted network with many knobs, nodes, and layers. Most of the above cells are switches in the down position. In each row the single switch in the up position is the word the row represents. The property of having one switch on in each row is where the name “one hot encoding” comes from. Current LLMs have an input switch array representing thousands of words over a dictionary of tens of thousands of tokens.

The above representation is a minimal idea that works. It is incredibly inefficient- taking tens of thousands of input switches (or voltages) to represent a single word. It is fairly rigid as the exact positions of the switches are used to encode the words, disallowing ideas such as using different sets of switches for different regions of the input document. And it understands nothing: two rows are either identical (have the same switch up) or disagree in exactly two columns (there is at this point no notion of similarity or synonyms).

Clever un-censorship

Now we are ready for the clever bit. I first saw this in an important research paper introducing a neat text embedding (defined later) called word2vec. The clever idea is the following.

Take our switch array and turn off all of the switches in one row. In this case we have suppressed the fourth row which used to encode “fox.” We wil treat the voltages implied by the new switch array as a single training input.

a

brown

dog

dug

earnest

fox

jumped

lazy

over

quick

red

slow

the

under

the a brown dog dug earnest fox jumped lazy over quick red slow the under
quick a brown dog dug earnest fox jumped lazy over quick red slow the under
brown a brown dog dug earnest fox jumped lazy over quick red slow the under
? a brown dog dug earnest fox jumped lazy over quick red slow the under
jumped a brown dog dug earnest fox jumped lazy over quick red slow the under
over a brown dog dug earnest fox jumped lazy over quick red slow the under
the a brown dog dug earnest fox jumped lazy over quick red slow the under
lazy a brown dog dug earnest fox jumped lazy over quick red slow the under
dog a brown dog dug earnest fox jumped lazy over quick red slow the under

The weighted network will produce an output of meter readings as a function of the input (given above) and the positions of the weight/parameter knots. Here is one possible output.

a

brown

dog

dug

earnest

fox

jumped

lazy

over

quick

red

slow

the

under

a brown dog dug earnest fox jumped lazy over quick red slow the under

Now we construct a single row switch array encoding the missing word (in this case “fox”). Treat the voltages from this switch array as the desired training output.

a

brown

dog

dug

earnest

fox

jumped

lazy

over

quick

red

slow

the

under

a brown dog dug earnest fox jumped lazy over quick red slow the under

Training is just jittering the weight/parameter knobs a small bit so that the result meter needles for this input are closer to zero in the wrong words (down switch positions) and 10 volts in the target word (up switch position). The more serious term for this is stochastic gradient descent. What is amazing is small improvements can accumulate, instead of canceling each other out as we move from training example to example. The training procedure embodies the lesson of the LLM methodology: harvest an unreasonable number of small improvements to yield an approximation of a seemingly impossible desired outcome.

This one sentence could in fact give us 9 training examples- as we cycle through which word-position we wish to un-censor. We use these training examples and many others to train up a weighted network that simulates un-censoring a single word out of sentences! Obviously the simulation can’t be perfect (censorship loses information), but the weighted network settings route signal to plausible replacement word positions. Training is about appropriateness (using only sensible utterances as training data, so there are word relations to learn) and scale (having a lot of training data, for commercial LLMs: most of the web, must help/discussion forums, most social media, most books and periodicals, most technical and scientific papers).

After training is finished we use the weighted network as before on inputs. However we decode these outputs by picking a highest indicating meter as the plausible answer (in this case the 6th meter, which is +10v at “fox”) and get the following one-hot style result.

a

brown

dog

dug

earnest

fox

jumped

lazy

over

quick

red

slow

the

under

a brown dog dug earnest fox jumped lazy over quick red slow the under

Note: the word2vec paper popularized an additional concept of a semantic embedding. One of the layers of the word2vec weighted network was restricted to be only 300 nodes, much smaller than the input or output layers. This “constriction” layer tends to result in trained networks where similar meaning words yield similar voltage patters at the intermediate layer (and different meaning words induce very different voltage patterns). This is in contrast to the original one-hot encoding where different words always disagree in exactly two positions (so there is no useful notion of similar or dissimilar). Embeddings went on to be a core idea and produce additional products such as semantic databases.

The clever training method has given us a new encoding of words- where words that can be used in similar places tend to get similar numeric representations. We are now ready for the last of the big ideas: attention heads.

Attention heads

In my opinion the final component that explains current LLM performance is “attention”, a term defined in the paper “Attention is All You Need.” Attention is a bit anthropomorphic, so I will call them attention heads; think of them as signal routers, position pointers, or text bookmarks.

Current LLMs treat all of your input instructions, the input text, and currently generated output text as encoded inputs as we described above. At first there is no output, so only the system prompts and user inputs are encoded. Then highest LLM-scored word is chosen as the next output word. This process is then repeated by the LLM operator to generate the output text one word at a time in order. By our analogy we are pretending there is a pre-existing plausible output but it has been censored or hidden and the LLM operator is un-censoring an approximation of the output one word at a time. This repetition isn’t part of the LLM, it is supplied by the operator serving the LLM results.

“Attention heads” are just bookmarks that point to positions in the input (and also partially generated output) and also to earlier sections of the weighted network. The positions of the heads are a function of the inputs, so they can appear to move between re-applications of the LLM. This allows effects such as generating words or tokens to complete a sentence (by having a head point into the uncompleted sentence until an ending token is generated), starting or stopping a paragraph, using the same name, and many other seemingly meaningful long-range text interactions. It even can allow semantic effects such as making a point only once, by not pointing the attention head at a given input question if there is what appears to be a related answer already in the partial output. This is likely why LLMs appear to follow instructions- they leave attention heads in the instruction region of the text input.

It is likely the amount of training material and training time is rises very fast in the number of attention heads (probably even exponentially fast). So attention heads are likely expensive even for the rich. I believe they will be a limiting feature of LLM output for a while.

Conclusions/Speculation

I hope you now can envision LLMs as transformations on text realized as a very brutal encoding of input and partial-output text into a flow of numeric signals. The LLM server iterates a “plausible next word” process to generate a sequence of output tokens, as the LLM itself doesn’t implement repetition or iteration. The bookmarks or attention heads help prevent this process from drifting too fast to text unrelated to the original inputs.

I believe features of LLM text (good, bad, and amazing) can be explained in terms of word clustering, approximate un-censorship, and bookmarks. That is we are not forced to explain observed texts in terms of goals, desire, intent, and memory.

LLMs are largely a triumph of scale (size of net, number of knobs/parameters, size and diversity of training data). The LLM weighted networks are of previously unimaginable size. The LLM training corpus can be effectively many times larger than the sum of all written text, as the un-censor training procedure can build many examples from a given text. And LLMs can show amazing results on tasks that don’t need too many attention heads such as specializing or translating a description of a mathematical or engineering technique from the LLM training data into a specific problem application at query time. However, current LLMs have poor performance on seemingly simple tasks that burn attention heads: such as my example of failing at uniquely sorting words.

Unless you are taking the trouble to run a local model, you never directly observe LLM behavior independent of the infrastructure and staff of the service provider. It is hard to tell what is LLM behavior, and what is cached-results or additional custom tools or services. For example LLMs themselves can’t repeat or iterate, however LLM service results are usually the result of iterating “generate next plausible word.” What one is seeing is the result of the LLM plus the service stack.

I’d invite you to try to apply the above (simplified, so not fully realistic) model of LLMs to try and think on a number of scenarios. What are your opinions on the likely outcomes and results of the following thought experiments>

  • We take a known mathematical fact (say the Pythagorean theorem), and ask the LLM for a proof.
  • We take a pre-existing published math paper, present the first half to the LLM and ask that it complete the paper.
  • We take a pre-existing published empirical chemistry paper (concentrating on lab work and results), present the first half to the LLM and ask that it complete the paper.
  • We take a new non-published empirical chemistry paper (concentrating on lab work and results), present the first half to the LLM and ask that it complete the paper.

Your answer to the last should vary depending if you treat the LLM as an approximate text processor, or as having a tiny chemist and lab trapped somewhere in its clockworks.

Appendices

GPT-4

GPT-4 is partially documented here. In particular (quotes from article):

  • “Transformer-based model pre-trained to predict the next token in a document.”
  • “GPT models are often trained in two stages. First, they are trained, using a large dataset of text from the Internet, to predict the next word. The models are then fine-tuned with additional data, using an algorithm called reinforcement learning from human feedback (RLHF), to produce outputs that are preferred by human labelers.”

Note we don’t comment on the full complexity of transformers, just the attention head feature.

On Hallucination

From the GPT-4 Technical Report.

GPT-4 has the tendency to “hallucinate,” i.e. “produce content that is nonsensical or untruthful in relation to certain sources.”

I think of this as anthropomorphizing, but the online dictionaries don’t seem to support me:

hallucinate (third-person singular simple present hallucinates, present participle hallucinating, simple past and past participle hallucinated)

  • (ambitransitive) To seem to perceive things (with one or more of one’s senses) which are not really present; to have visions; to experience a hallucination.
    Synonyms: imagine, see things
  • (artificial intelligence, of a model) To produce information that is not supported by the model’s training data.

Joking aside: for a non-technical audience “hallucinate” evokes perception and mental state, not a mismatch from training data. The LLM output was always a construct or fabrication, even when it matches the truth.

On Recursive Self-Improvement

Investors have been hoping for recursive self-improvement where it turns out one of the things LLMs are good at is suggesting game changing improvements to LLMs (and then trigger some sort of technological singularity). Things are in fact moving very fast in the LLM space. So there are some issues in working out if unfounded confident-sounding advice from an LLM is the best steering for a long and expensive training cycle. Optimizing in the presence of delayed feedback is notoriously treacherous.

Note on Experiments

Note both the word2vec and attention papers are very good experiments in that they deliberately use overly simplified complementary tools and procedures to show the claimed positive effects are from the claimed technology. So not only are the results reproducible, one can do better by swapping in better complementary tools (such as tokenizers, embedding choices, and so on).

Things LLMs should be bad at

Under our mental model LLMs should produce bad results when there is a dominant plausible wrong answer and when there are many relations between bits of the answer to maintain. The ideas being if no answer is plausible the LLM result will likely be equivocal and unconvincing. When there are a lot of relations to maintain in the answer (such as puzzle conditions) then attention heads are used up mapping marking relations between parts of the answer text, moving them off relations to the input text and so-called instructions and question. Another attack is to have an input text that superficially looks like a common puzzle, this way the LLM output tends to be aligned to the related answer.

Some failing examples include:

  • Failing to uniquely sort words. This is exploiting the presumably limited number of attention heads.
  • Failing on the “surprise the doctor is your Mother, not a man!” puzzle. This is the result essentially being a copy of the answer to a similar puzzle, not the one asked. Note the claimed “reasoning” is better described as “initial tokens.” Notice the so-called reason trace does start with text close to the question and with text at the correct answer, however it is full of weird non-sequitur stops and starts. It is text that plausibly looks like reasoning, probably not reasoning. The ideas of additional state heads and state-reprocessing are in fact good, they just do not necessarily work in the way the author or even I think.
  • The “should I walk to the carwash” example. Trick questions often exploit missing context or corner-cases of reasoning. This is not in fact a trick question: for a human going to the car wash almost implies it is to get the car washed. Yes it could be to buy an air-freshener or pick up an already committed car, but those are the exceptional cases. That the LLM suggesting walking is evidence against it using non-textual semantics.

The question isn’t: can we make silly examples LLMs get wrong. It is: are the equivalents of these problems lurking in our important project we delegated to the LLMs?

Image credits

https://commons.wikimedia.org/wiki/File:Mcintosh_MC275_european_version.jpg#/media/File:Mcintosh-MC275-glow.jpg ,
https://hackaday.com/wp-content/uploads/2022/09/AMSAI.png , and the author.

Read the whole story
mrmarchant
19 hours ago
reply
Share this story
Delete

The Cables that Connect the World

1 Share

At the northeastern edge of La Línea de la Concepción, on a scrubby Mediterranean beach called El Burgo–Torrenueva, there is an old battlement-tower, La Torre Nueva, and not much else. It was part of the system of coastal watchtowers during the 16th century that would defend the area against the incursion of the Barbary corsairs. The coordinates are 36°12′36″N, 5°19′27″W. Walk the tideline and you would never know that buried two metres beneath the sand, a fibre-optic cable comes out of the sea here and turns into the internet. It’s the start of a line that runs across the Strait of Gibraltar to Ceuta, on the African coast, and on toward two continents. Nearly everything you do online that crosses an ocean passes through a cable like this, ending, in most cases, underneath a similarly unremarkable patch of coast.

Note: Ceuta is an interesting place by itself, that has recently gained some attention and that would also make for an interesting write-up of its own. However, the tl;dr is that it is an autonomous Spanish city of some 85,000 people sitting on the North African coast, bordering Morocco, which means the European Union has one of its very few land borders with the African continent running straight through a peninsula most people could probably not even point to on a map.

It has been held by the Spanish crown since 1668, it had been Portuguese before that, and Morocco seemingly never stopped claiming it. For our purposes, though, what matters is that the small enclave, until very recently, hung off the mainland’s network by a single ageing link.

When we talk about the internet we do so as if it were air. Ambient, ownerless, and everywhere. In reality, however, it is the exact opposite, because international data doesn’t (normally) travel by, let’s say, satellite, despite what most people might assume. It travels through roughly 1.5 million kilometres of very real (and very owned) fibre-optic cable lying on the seabed, surfacing at a small number of carefully chosen landing points.

For these landing points you normally need a gently sloping seabed, mild currents, and little marine traffic, so that anchors and trawlers don’t sever the line. Suitable spots are scarce enough that the same beach usually becomes the shared landfall for several cable systems at once.

Cables? What cables?

Unlike what you might be thinking of at first, submarine cables aren’t your run-of-the-mill Ethernet or fibre cable. The hardware that does the heavy lifting out in the deep ocean is about as thick as a garden hose with roughly 25mm across and weighing in at around 1.4 tonnes for every kilometre. The part that carries your data is a small bundle of glass fibres, each one around the same thickness as human hair, sitting in the very middle.

Everything else wrapped around those fibres is there to keep them alive in a deeply hostile environment. Working outward from the core, the fibres sit in a water-blocking gel inside a thin copper or aluminium tube, which is sheathed in polycarbonate, then an aluminium water barrier, then a layer of stranded steel wires that give the cable its tensile strength, then a wrap of mylar tape, and finally an outer skin of polyethylene. The copper is for power, because the cable doubles as a very long extension lead, which we will get to in a moment. Closer to shore, where trawlers and anchors roam, the whole thing gets one or two further jackets of galvanised steel armour wire, swelling it to 50mm or more in diameter and several times the weight. Hence, the cable that surfaces on our Spanish beach is buried a couple of metres down and not simply left lying on the sand.

The reason a copper conductor runs the entire length is that light, no matter how pure the glass, slowly fades as it travels, and so every 50 to 80 kilometres the cable is interrupted by a repeater, which is an optical amplifier that boosts the signal back up before passing it along. Each repeater needs electricity, and because the fish sadly still didn’t manage to install power sockets on the ocean floor, the shore stations at either end have to feed a direct current of anywhere between 3,000 and 15,000 volts down that copper core, to literally power the cable from both ends at once.

On top of the amplification, modern systems lean on a stack of clever tricks to keep the signal intelligible across thousands of kilometres of glass, including wavelength-division multiplexing to cram many separate colours of light down a single fibre, coherent detection to read them back out, and forward error correction to repair whatever gets garbled along the way.

Length, then, is mostly a question of power and amplification rather than of the glass itself. Shorter hops can dispense with repeaters entirely, hence an unrepeatered span will happily run to around 250 kilometres on amplifiers at each end alone, which is roughly the length of the line we started this post with. At the other extreme, a single system can stretch across an ocean, and the longest of them, like the 2Africa cable encircling the continent it is named after, run to tens of thousands of kilometres.

Who is laying cables?

The actual manufacturing and laying of these cables is, perhaps a little surprising for something the entire global economy rests on, the business of only a small handful of companies. The bulk of the world’s submarine cable is built and installed by just four suppliers, namely the American SubCom, the French Alcatel Submarine Networks, the Japanese NEC, and the Chinese HMN Technologies. They own and operate the specialised fleet of cable-laying ships, which aren’t exactly the kind of boat you would recognise from a harbour, but more like a purpose-built vessel carrying thousands of kilometres of cable coiled in enormous tanks below deck, rolling it out over the stern at a steady walking pace as they crawl across the ocean.

Deploying a new system is a multi-year effort that begins long before any ship leaves port. First somebody, these days increasingly a content giant rather than a phone company, decides a route is worth having and assembles the money for it, either alone or as a consortium of several owners sharing the bill. Then comes a marine survey, in which a ship maps the intended path along the seabed to find the gentlest, safest route around wrecks, trenches, and other people’s cables, followed by the permitting, which is the paperwork of securing landing rights and concessions from every jurisdiction the cable so much as touches. As we are about to see on the Spanish beach, this can generate a remarkable quantity of bureaucracy.

Only once all that is settled does the cable get manufactured to length, loaded onto the ship, and laid, with the vessel simply lowering it onto the seabed in deep water and a sea plough burying it a metre or two beneath the sediment closer to shore, where the danger from fishing and anchors is greatest. A working ship covers somewhere in the region of 100 to 200 kilometres a day, so an ocean crossing takes several weeks at sea.

A transatlantic system running some 7,000 kilometres typically costs in the order of 250 million USD, while a longer trans-Pacific route can easily climb towards 400 million, and the cable itself runs anywhere from roughly 6,000 to 20,000 dollars per kilometre, depending on how many fibre pairs it carries and how heavily it is armoured. Keep in mind that the spending does not stop once the cable is lit, because a submarine cable has a design life of only around 20 to 25 years and on top of that there are somewhere between 150 and 200 faults occurring across the world’s cables in a typical year. The overwhelming majority of them are not caused by sabotage or sharks, but by the combination of fishing gear and dragged ship anchors. Each break has to be mended by sending out one of a small number of dedicated repair ships, that are on permanent standby under regional maintenance agreements, to grapple the cable up off the seabed, haul both severed ends to the surface, splice them back together, and lower the repaired thing back down. This is slow and weather-dependent work that is quite expensive.

Who owns the cables?

With the data provided by TeleGeography’s Submarine Cable Map I have put together a list of the (co-)owners of undersea cables and sorted it by the number of cables each individual company has a stake in. The full dataset runs to some 473 distinct owners, the overwhelming majority of which are obscure national and regional carriers you will never have heard of, so rather than just dumping the entire list here, I limited it to the hundred most prolific (co-)owners:

(Co-)Owner # of Cables
Google 34
Orange 29
BT 22
Sparkle 22
Vodafone 20
Meta 19
Telekom Malaysia 19
Liberty Networks 18
Singtel 18
Tata Communications 18
Telkom Indonesia 18
AT&T 17
Telefonica 17
China Telecom 16
Chunghwa Telecom 16
NTT 15
Telstra 15
Telecom Egypt 14
XLSmart 14
China Mobile 13
China Unicom 13
GlobalConnect 13
Arelion 12
EXA Infrastructure 12
Verizon 12
e& 11
KT 10
Moratelindo 10
Softbank 10
Telxius 10
Altice Portugal 9
center3 9
KDDI 9
National Telecom 9
PCCW 9
Bharti Airtel 8
Globe Telecom 8
PLDT 8
Telin 8
Zain Omantel International 8
Djibouti Telecom 7
GCI Communication Corp 7
Indosat Ooredoo 7
Mauritius Telecom 7
Microsoft 7
Rostelecom 7
Colt 6
Entidade Administradora da Faixa (EAF) 6
Hawaiian Telcom 6
Setar 6
TDC Group 6
TIME dotCom 6
Triasmitra 6
Viettel Corporation 6
Zayo 6
Amazon Web Services 5
Bayobab 5
Camtel 5
Cyta 5
Dhiraagu 5
euNetworks 5
FLAG 5
Grid Telecom 5
Liquid Intelligent Technologies 5
Maroc Telecom 5
Ooredoo 5
OPT 5
OPT French Polynesia 5
Sri Lanka Telecom 5
Telkom South Africa 5
América Móvil (Claro) 4
Antel Uruguay 4
Bell Canada 4
Bharat Sanchar Nigam Ltd. (BSNL) 4
Bulk Infrastructure 4
Lebanese Ministry of Telecommunications 4
Libya International Telecommunications Company 4
Mobily 4
Okinawa Prefecture 4
Ooredoo Maldives 4
Pakistan Telecommunications Company Ltd. 4
Reliance Jio Infocomm 4
Starhub 4
SUBCO 4
Syrian Telecommunications Establishment 4
Tampnet 4
TeleYemen 4
Unified National Networks (UNN) 4
VNPT International 4
Vocus Communications 4
Whidbey Telecom 4
Algerie Telecom 3
Angola Cables 3
Australia’s Academic and Research Network (AARNET) 3
Bahamas Telecommunications Company 3
Bandwidth and Cloud Services (BCS) 3
Bangladesh Submarine Cable Company Limited (BSCCL) 3
BW Digital 3
Cabo Verde Telecom (CVT) 3
CANTV 3

Note: These figures are derived from the public Submarine Cable Map data, counting both, systems already in service, and those still planned or under construction (603 of the former, 91 of the latter, at the time of writing). The owners field is free-form text, so a few owners turn up under more than one spelling, and I had to do a little manual untangling of company names.

What jumps out, at least to me, is the name sitting right at the top. For most of the history of this infrastructure the owners were telephone companies, the _BT_s and _AT&T_s and _NTT_s of the world, laying cables to carry one another’s calls and, later, traffic. Google now has a stake in more submarine cables than any traditional carrier on the planet, with Meta not far behind, and Microsoft and Amazon both slowly accumulating their own share. The companies that fill those cables with traffic have, over the past decade or so, decided that they would rather own the pipes than rent them.

The other thing the numbers tell you is just how long the tail is. Of those 473 owners, some 260 appear on exactly one cable, and more than 340 of them, north of seventy percent, on no more than two. These are the world’s national telecoms, each one buying a slice of the handful of consortium cables that happen to land on its particular stretch of coast, which is also why so many of the big international systems list a dozen or more co-owners apiece. The internet, seen from this angle, is less of a single network and more of a mix of local operators, all chipping in for a share of the same few very expensive ropes across the ocean.

Going back to the beach in Spain

To see what it looks like where the cable actually meets the land, let’s head back to that beach in La Línea.

The cable that surfaces there is called Dos Continentes, it belongs to GTD, a Chilean telecoms group, and it’s a relatively small regional system consisting of two armoured fibre cables looping across the Strait of Gibraltar to Ceuta, the Spanish enclave on the African coast that depended on a single ageing link before this one was built.

I went looking for exactly where it comes ashore, and the paper trail gives an idea about how invisible this infrastructure actually is. The cable lands in Spain, but the public Spanish government map of coastal concessions doesn’t seem to show it, because it looks like coastal permits in Andalusia are devolved to the regional government. The landfall instead shows in a regional registry, in a signed resolution buried under an expediente number. That document pinpoints where the cable enters the public maritime domain, at grid reference X=290,935, Y=4,009,603, just seaward of the beach manhole. The cable then runs inland, buried as the permit insists (“no exterior element above ground level”) to what is presumably a network node, where traffic is fed into GTD’s pre-existing terrestrial dark-fibre network, from where it’ll eventually travel to one of the actual GTD data centres in Madrid, Barcelona, Bilbao/Sopelana, and Sevilla.

On its way out to sea it crosses three older cables already lying on the seabed, namely Europe India Gateway, ATLAS, and FLAG. As can be seen (or, well, actually not) even an empty-looking patch of water off a Spanish beach is layered with other people’s infrastructure.

Note: When GTD applied, it seems that the town council of La Línea formally objected and asked them to drop the project. The cable, the council said, cut straight through the main local fishing ground, “splitting it literally in two”, threatening the small shellfish and trasmallo boats that work those waters, and a protected limpet that lives on the rocks, in a town whose fleet was already squeezed by run-ins with Gibraltar over fishing rights. However, they were overruled and the concession was granted anyway, with mitigation conditions attached, for an initial fifteen years.

The Dos Continentes cable (Segment I, La Línea - Ceuta Sur ramal), owned by GTD Cableado de Redes Inteligentes, S.L.U., the Spanish arm of the Chilean GTD group, has a total length of ~105 km and is in service since 2020 under the signed concession resolution from the Junta de Andalucía (Dirección General de Calidad Ambiental y Cambio Climático), expediente CNC02/19/CA/0009, dated 14 January 2020.

The two key points, as given in the resolution’s coordinate table are:

Point UTM X UTM Y
Arqueta / beach manhole (BMH, in servidumbre zone) 290,929 4,009,602
Entrada en DPMT (cable crosses into public maritime domain) 290,935 4,009,603

Note: The resolution’s prose text gives a slightly different value that disagrees with its own table by approximately 140m.

To convert the UTM coordinates I used the official Instituto Geográfico Nacional (IGN) Calculadora Geodésica with the following settings:

  • Transformation type: Transformación de Datum
  • Reference system: ETRS89
  • Input coordinates: UTM
  • Huso (zone): 30

ETRS89 and WGS84 differ by only centimetres in practice, so the resulting coordinates (WGS84-equivalent) can be dropped straight into any consumer map or GPS app:

Point Lat/long (DMS) Decimal Map links
Beach manhole (the buried structure, navigate here to stand on the spot) 36° 12′ 31.25″ N, 5° 19′ 32.41″ W 36.208681, −5.325669 Google Maps · OpenStreetMap
DPMT entry point (waterline crossing, ~6 m seaward of the manhole) 36° 12′ 31.29″ N, 5° 19′ 32.17″ W 36.208692, −5.325603 Google Maps · OpenStreetMap

Both points sit on Playa de El Burgo–Torrenueva, beside the Punta de Torrenueva tower, at the northeastern (Levante / Mediterranean-facing) edge of La Línea de la Concepción, against the municipal boundary. The resolution describes the route as passing “muy cerca de la torre-faro existente en la Punta de Torre Nueva”.

As you can see, however, you see nothing. :-) The permit requires the whole installation to be subterranean (“no exterior element above ground level: No manholes, splices, connections or terminals.”), hence you can stand exactly on the landfall, but it’s a point in the sand by a tower, and not a structure. On the afternoon I was there, a couple of dozen people were spread out on that stretch of sand under parasols, probably not even knowing that somewhere underneath them the link that carries an entire enclave’s traffic to another continent came out of the sea.

It is interesting to see that what has changed most over the past decade isn’t the technology itself, but who pays for it. For a century these systems were built by carriers selling capacity to one another, which made the network something close to a shared utility with many owners. Today, however, the largest (co-)owner of submarine cable on the planet is an advertising company. It probably makes sense in their position, however it is a change in how the network is governed, and, more importantly, it seems to have happened almost entirely out of public view, which is worrying.

If you live anywhere near a coast, there is a decent chance one of these things lands within driving distance of you, and the TeleGeography map will get you to roughly the right bay. Getting from there to the actual patch of sand takes some amount of digging through concession resolutions, planning registers, environmental reports, and sometimes the local newspaper archive. It took me an evening of reading to narrow it down, but I can recommend to do this exercise if you’re curious about the world that you’re living in and, more importantly, the hidden infrastructure surrounding you.

PS: Maybe we picked the wrong word and should have called it the trench rather than the cloud?

Read the whole story
mrmarchant
1 day ago
reply
Share this story
Delete

How Much of the Internet Is Written With AI?

1 Share
In a random sample of 10,000 webpages collected in July 2026, one-in-ten show signs of being written or substantially edited by AI.
Read the whole story
mrmarchant
1 day ago
reply
Share this story
Delete
Next Page of Stories