October 2026

Saturday, October 10, 2026

Writing

Agents drop spaces when talking to each other

"...actualall-datesretained withinexisting62snapshot..." says Sol

Read here: Agents drop spaces when talking to each other

One finaltarget improvement realmail preparemodel saw50irrelevant M1purchases and uncertain; root scheduledmonthly/pin now filters candidates sameObligation+negativeamount (globalcross-source recurrence ambiguity still intact). This is financialdirectionidentity pruning not datecycle cutoff, actualall-datesretained withinexisting62snapshot. Newcycle tests same. No livewrite. Readonlyrealpreparepending.

  • GPT 6.1 Sol when speaking to another agent.
Martin Emde

Sunday, October 4, 2026

Writing

reverse.horse - The feeble human brain powered Jev API

An attempt to understand and explain the new AI model Jev, by using your feeble human brain as the model.

Read here: reverse.horse - The feeble human brain powered Jev API

Reverse.horse is an attempt to understand and explain the new AI model Jev, by using your feeble human brain as the model. Place yourself in-request and respond as best as you can to incoming questions.

If you want to jump right in, then go try it at https://reverse.horse and come back here for the reintegration with humanity. The horse will put you through your training on how to be the AI model and then cut you loose to do the job of the AI model. You may come back with a new appreciation of what these models are doing.

What is Jev?

Jev is a new AI model that has taken the internet by storm in recent weeks. The concept is based on the idea that if we make “good enough” intelligence cheap and fast, it will explode in usefulness, expanding the use of intelligence into more places than we ever thought practical.

Jev is a different direction than most LLMs have taken. ChatGPT and Claude try to be ever bigger and smarter to take over for your brain, while Jev asks only for ever-more-decomposed questions. You send Jev state and questions and you get the answers to those questions very quickly and cheaply. I picked this model to clone because it impressed me. It opened up new ideas for working with intelligence, and, like many great ideas, it seems so simple after the fact.

The API route systemone (that’s “System One”) implies the behavior of the model. System One comes from Daniel Kahneman’s Thinking, Fast and Slow, which defines it as quick, gut-reaction thinking. System Two is what almost all modern LLMs are doing: Reasoning…

The name “Jev” comes from Jevons Paradox, which has been a buzzy term for a characteristic of demand elasticity: when you make consumption of a resource more efficient, sometimes you use more of that resource than you did before. Cheaper electricity makes people apply electricity in places where it was previously too expensive to do so.

Why?

Because! Reverse.horse is something between a tutorial and an art project.

Reverse.horse is actually a functioning API that matches the Jev /v1/systemone format. You could, in truth, hook this up in place of your real usage of Jev and get terrible, slow, human powered answers (if anyone is listening). The questions you ask are presented to anyone that is connected, and the results become the response from the API. Send real questions to reverse.horse and any human connected can scramble to respond before the 30 second response timeout.

It’s also a tutorial intended to help you understand exactly what Jev is and what the 3 question types mean. By using your human brain to answer in place of the model, you get an intuitive sense of what Jev is doing.

It becomes immediately apparent that you suck at this. Jev can usually answer in under 1/5th of a second, network and all, and it can do thousands of questions like this in the time it takes you to answer a single one. You may even notice that for every answer that you carefully answer correctly, Jev is within a few points of what you answered. Remarkable, isn’t it?

Why “Reverse Horse”?

The name comes from Cory Doctorow’s reverse centaur, a situation created by AI that places the human in the position of being a meat body that does what the AI tells it. Instead of a good centaur, a human brain controlling a powerful horse body, you have the opposite: an “unthinking” horse brain controlling the feeble human body.

The irony here is that I’ve inverted it. By using your feeble human brain to answer as the “unthinking” AI, I’m pointing out just how slow you are at it, and how impressively fast and accurate Jev is. When I started the project, the example questions were actually too complex to answer within a reasonable timeout, so I had to break them down. Feeble human brain!

Why the 30 second timeout?

It’s kind of stressful, right? It’s there for 2 reasons:

  1. I’m trying to make you give your gut reaction to the question. The timer makes it clear just how slow we are at this. 30 seconds is 150 times longer than Jev usually takes to respond.
  2. This is a real, if impractical, API. You can actually send a request to it with curl. 30 seconds seemed like the longest reasonable timeout for a held-open request. I’m waiting for the moment when this 30 second timeout causes the poor server to get overloaded.

How does reverse.horse work?

In theory, everyone answering questions is connected to a web socket and will be asked live the questions that arrive at the API. Their answers are assembled into the response. That’s the goal, and it works across a small group, but it’s almost certainly bound to break if too many people visit. I look forward to finding out the strange new ways this can break.

How did I not know about the .horse TLD!?

Welcome to a whole new world of domain ideas! The domain reverse.horse is me stupidly spending $25 on a joke. Secretly, I’ve been chomping at the bit to have a good reason to buy a .horse domain for years. Every new idea I have gets a run through “could this be a .horse domain?” I already have another project planned that may end up on .horse. It’s my favorite silly TLD because it begs so many questions: Why horse? Why aren’t there other animal TLDs? Where’s the demand for this TLD coming from? How has this not created a new economic boom in horse related websites?

Enough silliness!

Please go try it out and let me know what you think.

Martin Emde

Saturday, October 3, 2026

Writing

OpenAI Is Winning – Does It Matter?

It's all based on a gut feeling and a couple of questionable graphs.

Read here: OpenAI Is Winning – Does It Matter?

OpenAI has the best coding models right now, but there’s a problem: These frontier labs are losing money on me. In fact, they’re losing a lot of money on everyone and their products seem practically designed to bypass any possible moat. The promise of AI: custom software and easier-to-use computers is directly at odds with vendor lock-in. I’m concerned about what that means for consumers, companies, and the US economy.

I write a lot of code with AI agents, both through my job on the AI Developer Tools team, and at home. I like OpenAI’s GPT Astra and GPT Sol better than Anthropic’s Claude, but it’s all based on vibes and a couple of questionable graphs. I was Team Claude for a while, but after switching to codex, I no longer feel attached to any provider. This is good for me but bad for the business models of expensive “luxury” token providers like Anthropic and OpenAI.

What kept me with Claude at first was comfort with the model: I know how Claude codes. But the questionable graphs kept coming. “Try GPT Sol,” the graphs said. I dabbled with a $20 subscription and many more tokens fit into the plan than I expected. Adaptation was easy. I stuck with the GPT family and only dabble in Claude (which remains better at chat).

Anthropic held the cutting edge from November 2025, often called “the inflection point”. OpenAI beat them at best coding agent per dollar with GPT 5.6 Sol and expanded their lead up through GPT 6.1 Sol.

Claude Fable, on the other hand, reinforces the vibes-beat-quality story and the questionable graphs back me up. I was left with a bad taste because of the price, the overzealous security safeguards, and the spat with the government regarding export controls. It’s become an expensive pariah for me. No one ever got fired for not using Fable.

Both subscriptions are vastly cheaper than the raw API price. Compare price per million tokens, or price per task, and the advantage over open weight models evaporates. Anthropic and OpenAI are effectively paying me to stay on their subscriptions. They have no moat and the frontier labs know it. The choice between models is all vibes and, ironically, the products they sell are built to help you switch. They’re built to do everything.

The frontier labs lack pricing power over individuals. An exodus of personal subscribers would spell doom for their narrative. Consumer intolerance for restrictive policies has already made them back away from disabling AgentSDK and cli usage, and the encroachment of open weight models into the questionable graphs is continuing to exert pressure. They’ll keep subsidizing until they can’t…

…or don’t need to. “But what if there was a way for the AI companies to get government permission to violate antitrust law and cease to compete with one another, and secure a ban on the use of Chinese open weight models?” asked Cory Doctorow recently. Monopolies do make it difficult to find a better competitor. So far this seems unlikely.

For now, I will use the luxury tokens from the current best provider while they last, but my loyalty is gone. If I’m a representative power user, and if that’s the source of most consumer spend, then it could pose a tricky problem for AI companies trying to support their valuations. The commoditization of intelligence will continually force subsidization of loyalty in order to maintain the narrative that keeps their training runs funded.

Anthropic anticipated more than half a trillion dollars in spending in their potential upcoming IPO filing. It also notes that revenue is growing quarter over quarter, and that they would be profitable if you ignore how much it costs. Rising revenue may not even be enough. As the second derivative argues, “I do not need demand to fail. I need the rate of capex growth to flatten - and a structure this levered and dependent on perpetual acceleration breaks on the flattening alone.” Even with revenue growing, can these subsidies hold in the face of a product that has near zero lock-in and questionable consumer loyalty? They both have to and, I think, can’t afford to.

The math for enterprises is different. Most of them pay full API prices and the terms of their lock-in are different, often self-inflicted. I’d like to examine that soon.

Martin Emde
martinemde.com/2026/10 © 2025 Martin Emde v2026.7_