Dropstone is not a foundation model. It is a runtime that turns open-weight models into functional agents, and the version number tracks the integration cycle rather than the weights. Each cycle we evaluate the strongest open-weight models available and ship whichever wins. In 1.7 all three tiers changed: Fast moves to DeepSeek V4 Flash 0731, Pro to GLM-5.2, and Heavy to Kimi K3. The headline is that Kimi K3 scores 57.1 on the Artificial Analysis Intelligence Index, the highest of any open-weight model, and the first open model to clear GPT-5.5. It does not catch the closed frontier. Claude Opus 5 sits at 60.7. We will come back to that gap rather than skip past it. Heavy runs Kimi K3. Pro runs GLM-5.2. Fast runs DeepSeek V4 Flash 0731. The Pro change is the one most users will feel. GLM-5.2 was 1.6's Heavy tier: a 744-billion-parameter model, roughly 40 billion active per token, MIT licensed, 1M context. It is now the Pro default, which means the model that was gated to paid plans six weeks ago is available on every plan including free. Moving it down a tier is not a demotion of the model. It is a statement that what counted as our top tier last cycle is now our baseline. This is the part of the release we nearly under-reported, and it is the largest single move in it. DeepSeek V4 Flash 0731 shares its architecture and its price with the V4 Flash we shipped at 1.6: 284B total, 13B active, 1M context, same rate card. But it is a retrained model rather than a point revision, and Artificial Analysis measures it at 50.0 on the Intelligence Index against 40.3 for its predecessor. Ten points is a bigger move than anything else in this release, Heavy included. Three things follow. Fast is now 6 points ahead of DeepSeek V4 Pro, a model a tier above it in its own family, while activating 13B parameters per token. Fast is 1.1 points behind our own Pro tier, and both are free on every plan, so the capability floor of a free Dropstone account moved to within a point of its ceiling. And the top three open-weight models on the index are Dropstone Heavy, Pro and Fast: Kimi K3 at 57.1, GLM-5.2 at 51.1, and V4 Flash 0731 at 50.0. We did not arrange the lineup to produce that sentence. The three models we selected independently on capability are the three that occupy those positions. At 1.6 we described Fast as a latency tier and made no capability claim for it. That description is no longer accurate, and we would rather correct it than let it stand. Three closed models are ahead of the best open-weight model in the world. Claude Opus 5 leads Kimi K3 by 3.6 points on the aggregate index, Fable 5 by 2.8, and GPT-5.6 Sol by 1.8. The two agentic Elo measures repeat the pattern: GDPval-AA v2 puts K3 at 1668 against Opus 5 at 1861, and AA-Briefcase puts it at 1547 against 1720. The table below is the one we would least like to publish, so it is the one we lead the comparison with. Heavy does not win a single capability row on it. It wins the licensing row, and that is the row the rest of this post is about. What changed is the size of the premium. At 1.6 the equivalent gap was 8.8 points, and at that distance the frontier was a different class of tool. At 3.6 points it is a margin most workloads will not notice. The ones that will are identifiable in advance: long-horizon autonomous runs, where the Briefcase gap is proportionally wider than the aggregate index suggests and shows up as the frontier drifting less over many steps. If your work is dominated by multi-hour agent sessions, that is the difference to weigh. The gap narrowed by more than half in one cycle. It did not close, and we are not going to describe it as though it did. There is one independent measure where the open models come out ahead, and it is the one with developers in the loop rather than a scoring harness. On the Frontend Code Arena a developer is shown two anonymous outputs and picks the better one, without being told which model produced either. Across 1,757 valid votes announced on 16 July 2026, Kimi K3 ranked first at 1679, ahead of Claude Fable 5 at 1631 and GPT-5.6 Sol at 1618. GLM-5.2, our Pro tier, placed fourth at 1587, ahead of Claude Opus 4.8. K3 took six of the seven frontend categories, placing second only in game. One generation earlier, Kimi K2.6 sat eighteenth on this same board. Two qualifications, and we do not want them read as small print. Claude Opus 5 is not on this board. It released on 24 July, after the ranking was published, and its Frontend Code score has not been published yet. First place here means first against the field as published, not against the closed frontier as it stands today, and the standing should be re-checked once Opus 5 is scored. Separately, a single leaderboard measures a single thing, and this one measures frontend code. It is not evidence about coding in general, and we do not extend it to backend work, long-horizon agentic tasks, or the aggregate index, where the ordering above does not hold. With both of those attached, it is still the strongest independent result in the release, and the only one decided by developers looking at output rather than by a benchmark harness. 1.6 led with SWE-bench Pro. 1.7 cannot, and switching benchmarks between releases is exactly the move a dishonest report would make to manufacture a win. So here is the reasoning in full rather than in a footnote. Moonshot published no SWE-bench Pro result for Kimi K3. The coding numbers it did publish, including Terminal-Bench 88.3 and FrontierSWE 81.2, were produced on Moonshot's own KimiCode harness. When one model runs on its vendor's harness and another runs on Claude Code or Codex, the difference between the scores measures the two systems, not the two models. We do not chart those numbers anywhere and we do not cite them as evidence here. The Artificial Analysis Intelligence Index is not a weaker substitute. It is independently measured rather than vendor-reported. It is measured uniformly across every model we compare against, so there is no source asymmetry to disclose. And it covers all three of our tiers, where 1.6 could not place Pro or Fast on its hero chart at all. This is the first release in which the whole cascade sits on one axis. What the switch costs us is comparability with our own last report. You cannot line up GLM-5.2's 62.1 on SWE-bench Pro in the 1.6 report against Kimi K3's 57.1 here. They are different scales measuring different things. Where we compare across releases in this post, we measure both models on the current index rather than setting a 1.6 figure against a 1.7 one. That is the only honest way to run an axis change, and it costs us the clean year-over-year line a marketing post would want. DeepSeek published nine agent benchmarks alongside the 0731 retrain, including a DeepSWE score of 54.4 against 7.3 for the preview. We are not reporting those as measurements. They are vendor-run, DeepSWE used a harness that was not public at the time of writing, and no third party has reproduced them. The reason to be careful is visible in the one benchmark that appears on both lists. DeepSeek reports Terminal-Bench 2.1 at 82.7. Artificial Analysis measures 79. We use 79. Where a vendor number and an independent number disagree, the independent number is the number, and we would rather publish the smaller one and be right. The same rule retires Kimi K3's own coding benchmarks, which run on Moonshot's harness. We exclude them entirely rather than discount them, which means Heavy has no directly comparable coding-specific score this cycle. We would rather have that hole in the report than fill it with something we cannot stand behind. Fast and Pro are both text-only. Heavy has a native vision encoder but is gated to paid plans. So an image attached on any tier, on any plan including free, is routed to Kimi K2.7 Code, which has one. K2.7 Code was 1.6's Pro tier and is no longer selectable in 1.7, but it stays in the routing map for exactly this reason. That routing is what keeps screenshot-to-code working without a paid plan, and it is invisible: you attach an image to whichever tier you were already using and it is handled. It is also the one place where the model actually serving your request may not be the tier you chose, which we would rather state than let you discover. We report cost at the plan level only. Dropstone is Free, $20 a month for Pro, and $100 a month for Max. For comparison: Claude Pro and Max 5x are $20 and $100, ChatGPT Plus and Pro are $20 and $200, and Cursor Pro, Pro+ and Ultra are $20, $60 and $200. Metering is one weekly credit pool per account, shown as a single percentage with a reset countdown, rather than the rolling five-hour throttles the competition uses. We do not promise a fixed number of turns per week. Turn cost varies by task, and a fixed count is exactly the kind of number that is contestable the moment someone measures it. Two access facts about 1.7. Heavy stays gated to paid plans, because a single long-horizon Heavy task can consume a free account's entire weekly allowance, and gating it cleanly is better than letting a free user exhaust a week in one request. And Pro is now free-plan accessible and materially stronger than it was. GLM-5.2 on the free tier is the practical headline of this release for most users, more than anything that happened at the top of the lineup. The models swap every cycle. The runtime guarantees do not. Inference runs exclusively on US-hosted infrastructure. The open-weight models' own first-party APIs may be hosted outside the US; routing through Dropstone takes that off your compliance surface, and you configure nothing to get it. Every tool call, file edit, shell command and API call sits behind an approval gate before it executes, regardless of which model is behind the wheel. That boundary lives in the runtime, not the weights, which is what makes it hold when the weights change underneath it, as they just did on all three tiers at once. Changing two thirds of the lineup without touching the security posture is the whole reason the runtime is what carries the version number. We did not train the base models. Kimi K3, GLM-5.2 and DeepSeek V4 Flash 0731 are third-party open-weight models. We select, host and integrate them; we cannot audit their training data or weights, and we do not pretend to. We changed the primary benchmark axis this cycle, so the 1.7 numbers are not continuous with 1.6's SWE-bench Pro numbers and should not be compared against them. We think the new axis is better evidence and we have published our reasons, but a reader is entitled to treat a mid-stream axis change skeptically, and we would rather flag that than have it noticed. Kimi K3's weights are not public as of this writing. Moonshot has announced release for 27 July 2026 under a Modified MIT license, and until that lands K3 is an API-only model we are describing as open weight on the strength of an announcement. If the date slips we will correct this post rather than quietly restate it. Index scores move. Artificial Analysis re-runs models and revises its index; everything here is stamped v4.1, July 2026. And the Frontend Code Arena result predates Claude Opus 5, which has no score on it yet. If Opus 5 ranks above K3 when it is scored, that result supersedes ours and we will say so. Dropstone 1.7 is live today across Fast, Pro and Heavy. Fast and Pro are free on every plan; Heavy requires Pro or Max, because a single long-horizon Heavy task can consume a free account's entire weekly allowance. Full benchmarks, methodology and the complete limitations list are in the technical report at https://blankline.org/research/dropstone-1-7.