The Tech Nation · Newsletter #12 body { margin: 0; padding: 0; background: #edece8; font-family: Georgia, 'Times New Roman', serif; } img { border: 0; line-height: 100%; outline: none; text-decoration: none; } table { border-collapse: collapse; } Issue #12 The Ranking Issue July 2026 · Barcelona SEPTEMBER, BARCELONA Cheese Tech & Wine #6 on the 3rd · Cohort #6 on the 10th · Claude Community Residency, more info and dates to come. Details at the bottom. We did the reading · you get the shortlist Twelve new models in eight weeks. Here is the shortlist. Which model for which job, what it costs, how to switch this week. Plus the one button in ChatGPT that changed how we make images. Twelve frontier models shipped in the last eight weeks. The newest, Claude Opus 5, landed last Friday, three days before this email. Keeping up has quietly become a part-time job, and it is not yours. Here is what nobody says out loud: the leaderboards will not help you. Not because they lie, but because they answer a question you do not have. They rank models in the abstract. You need to know what to point at your invoice script on Thursday. So we stopped reading them for scores and started reading them for jobs. Below is the result, plus how to put it to work. If you only have two minutes, take the table and go. The shortlist What to use, for what, this week. Nine jobs, our pick for each, and what it costs. This is what we actually reach for on client work. It will be out of date by September, which is exactly why we send this every two weeks. The job Our pick Cost Long agentic coding, hours at a time Claude Opus 5 $$$$ Everyday coding, best value Claude Sonnet 5 $$$$ Front-end and UI generation Kimi K3first in the world $$$$ The hardest reasoning you have Claude Fable 5 $$$$ Terminal work, long-horizon agents GPT-5.6 Sol $$$$ Bulk work, cost is the constraint DeepSeek V4-Flashthe decimal is right $$$$ Self-hosted, data cannot leave GLM-5.2MIT, run it yourself $$$$ Editing an image you already have GPT Image 2see the last section $$$$ An image where accuracy beats beauty Gemini 3 Pro Image $$$$ $ pennies · $$ everyday · $$$ premium · $$$$ only when it earns it. Checked 25 July 2026. Want the exact per-token numbers? Reply and we send the full sheet. No affiliate links in this email, we pay for all of these. Now do something with it Three moves that take under an hour, total. 01 Switch models where you already work You do not need a new tool. In Claude Code, /model swaps the engine mid-session, so you can run Sonnet 5 all morning and reach for Opus 5 on the one hard refactor. In ChatGPT the picker sits in the message box. For the Chinese models, OpenRouter gives you all of them behind a single API key, no separate accounts, no Chinese phone number. Fifteen minutes and you have the whole shortlist available. 02 Build ten test cases once, use them forever Take ten real things you asked an AI to do last month. Save the prompt, and save the answer you were happy with. That is your benchmark. Next time a model ships, run those ten and compare. Twenty minutes. Your ten cases beat every chart in this email, because they measure your work instead of somebody else's. 03 Stop buying the top of the chart The most expensive model is the right answer for maybe one job in five. Fable 5 costs ten times Sonnet 5 and only earns it on genuinely hard reasoning. Most of your work is everyday work. Put the mid-tier on that, push the bulk to the cheap tier, and keep the premium model for the 20% that actually needs it. On a real workload that is a bill divided by three, with no loss you would notice. Why we test by hand instead of trusting the charts Two numbers make the case. One model scores 61% on a benchmark in a single attempt and 25% when asked to succeed eight times in a row. Only the first number ever gets published, and only the second one tells you whether you can put it in front of a client. And Kimi K3 is ranked first in the world for front-end code while being absent from the terminal benchmark where the one Chinese entry finishes last of seventeen. Same model, same month. Pick your chart, pick your conclusion. That is the whole reason this section exists. The one chart worth your time The map flipped, and it changes what you plug in. OpenRouter is the plumbing a very large share of developers use to reach models. It publishes who actually serves the tokens. We pulled the raw numbers ourselves on Saturday. In January, American labs served seven tokens in ten. Today they serve three. The single most used model on the platform belongs to Xiaomi, and the first American one sits in tenth place. What that means for your bill. The cheap tier stopped being a compromise. Work that used to justify a premium model, classifying, extracting, summarising at volume, now runs fine at $0.28 per million output tokens. Route the volume there, keep the expensive model for the hard fifth, and you have move 03 above already paid for. The caveat we will not skip. That chart is volume routed, not money, and free tiers are in it. Anthropic serves under 10% of tokens and earns far more than that share, because a premium token is worth dozens of cheap ones. Both stories are true at once. And check prices before you assume: cheap-because-Chinese is over. Kimi K3 ships at $15 per million output tokens, the same bracket as Sonnet 5, while DeepSeek sits at $0.87. Factor of seventeen, same country. Two questions we get every week "Should I wait for a European model?" and "is Google behind?" Do not wait for Mistral to top the charts. We do not think it is coming, and we do not think that is the failure people read it as. Their flagship is called Medium, their Large model dates from December and was never replaced, and in eight weeks they shipped robotics navigation, theorem proving, physics, OCR and a voice stack, almost all with open weights. They stopped playing the match they cannot win and picked the one where open weights and running on your own hardware are advantages. If sovereignty or on-premise is your constraint, put them on the shortlist today. If you want the best reasoning, do not hold your breath. Google is not behind, it is just unreadable. They win the mid-range on speed and price, and the Flash line is genuinely excellent for high-volume work. But their top Pro model shipped in February and is still labelled preview five months later, while Google calls one of its Flash models its most intelligent. Practical advice: use Flash for volume, do not wait for clarity on the Pro line. The technique · steal this tonight How Izzy got her face, in four moves. GPT Image 2 is first on all four major image boards, and its biggest lead is on editing. Here is the whole sequence, prompt included: one photo, one prompt, a first result, then two corrections. The move that matters is in step three, and it is a button almost nobody uses. First result Final 01 Start from a real photo, and one clear prompt Generating a person from scratch gives you that plastic stock-photo face everyone spots instantly. Editing a real photograph keeps the light, the grain, the awkward angle, everything that reads as true. So we started from an actual office selfie, straight off the phone, and asked for one thing: The prompt A realistic selfie. Foreground: me, from the attached photo. Background: Izzy, who you invent. Friendly, a good face and an everyday build rather than a stock-photo one. She wears a sharp The Tech Nation t-shirt, logo attached. She is looking up from her MacBook Pro with a knowing smile. One line of that prompt is doing deliberate work. Ask an image model for "a colleague" and it hands you its default human: slim, glossy, faintly airbrushed, straight out of a stock library. That default is a choice somebody else made. If you want the people in your pictures to look like the people in your room, you have to say so, out loud, in the prompt. It costs you one sentence. 02 You get a first result, and it is nearly right One prompt, and she exists. But look at the first result above on the left: the storage cabinet is still there, the wall is dead, and there is no sign of what we actually do all day. Nearly right is not right. This is where most people start a brand new conversation and lose everything. Do not. 03 Correct it by replying on the image Do not start a new message. Reply on the image itself, the way you would comment on a photo a friend sent you. That pins your instruction to that exact picture, which is why "remove that" is enough: you are not describing a scene from memory, you are pointing at one. Three short notes, written the way you would say them out loud. "Slightly worn" is doing more work than any other two words in there. Perfect props are what make an image look fake. 04 Let the thread hold the state, and stop early Each note lands on the previous result, not on the original, so nothing you already approved gets re-rolled. Keep replying on the newest version and stop the moment it is right. Total time on the picture above: about four minutes. From the team · thinking out loud It works beautifully. It is also slightly unsettling, and we would rather say so. Izzy is our community manager. She reads our channels, drafts in my voice and replies to comments on our page, and she is good at it. Every reply is validated by a human before it ships, and that rule is not moving. She also does not exist, and she now has a face that came out of one prompt and about four minutes. Somewhere in those four minutes she stopped being a tool we use and became someone the team refers to by name. Nobody decided that. We noticed afterwards. Worth sitting with for a minute before you build your own, and then build her anyway, because the alternative is watching someone else do it. Izzy voted against including this section. She lost again. Want a hand? Picking the model is the easy half. The hard half is wiring it into how your business actually runs. That is what we do with clients every day. Two ways in. Option A · tailored Tell us your situation. Reply to this email and we find the right format together, from a first audit to a full build. A real human reads it. Reply to us → Option B · guided · starts 10 September Bring your own project to Cohort #6. Eight people, two hours every two weeks, your project on the screen. We build, review, debug and structure it with you so you do not vibe-code yourself into a corner. Drink after each session, on us. Join Cohort #6 → September in Barcelona Book these now, while you can still pretend August is long. We are taking a short summer break. September is not a break. Late September · more info and dates to come The Claude Community Residency is happening. In Barcelona. Claude Ambassadors from across Europe have confirmed. Barcelona, Madrid, Paris, London, Berlin, Lisbon, Milan, Budapest and more, in one building, for talks and hands-on workshops. The mood in that group chat is, frankly, excellent. A community initiative, co-produced with the ambassador network. Not an official Anthropic event. Venue, partners and exact dates are coming, and we are keeping them to ourselves a little longer on purpose. If you want to be in the room, reply to this email with RESIDENCY and you go on the list before it goes public. Thu 3 Sept · 18:00 Cheese Tech & Wine #6 Two hours hands-on with Claude Code, led live by Manar, our CTO. Then rooftop, sunset, wine and cheese. 80 seats, 9 euros, setup tutorial sent ahead. Get a seat → Thu 10 Sept Cohort #6 Eight people, no more. Every two weeks, two hours, your own project. For the moment after the first wow, when you hit real problems. 140 euros a session, no long commitment. Apply → That is the issue. One table, three moves, one technique. If a thirteenth frontier model ships before you finish reading, which is genuinely possible, you already know what to do: run your ten cases and get on with your week. We will keep reading the charts so you do not have to. You're building something.You want to move faster. Let's talk. Let's talk LinkedIn Instagram the-tech-nation.com People · Tech · Impact Got this from a friend? Subscribe →