Close Menu
    Facebook X (Twitter) Instagram
    • Privacy Policy
    • Terms Of Service
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Facebook X (Twitter) Instagram
    Fintech Fetch
    • Home
    • Crypto News
      • Bitcoin
      • Ethereum
      • Altcoins
      • Blockchain
      • DeFi
    • AI News
    • Stock News
    • Learn
      • AI for Beginners
      • AI Tips
      • Make Money with AI
    • Reviews
    • Tools
      • Best AI Tools
      • Crypto Market Cap List
      • Stock Market Overview
      • Market Heatmap
    • Contact
    Fintech Fetch
    Home»AI News»Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks
    Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks
    AI News

    Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks

    August 15, 20264 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email
    kraken

    rewrite this content and keep HTML tags as is. This is content from rss feed and I don’t need their *Daily Debrief Newsletter*, their tags from bottom like this *Share this articleCategoriesTags*, Editorial Process section, phrases like *Featured image from Peakpx, chart from Tradingview.com*, SPECIAL OFFERS and similar sections – just remove such sections and save only article itself:

    Z.ai just released GLM-5.3. GLM-5.3 runs on the same 743B base model as GLM-5.2. Every reported gain comes from scaled post-training: more task environments, more environment types, longer training. The results land in two places. Coding jumps most on the longest-horizon benchmarks, with Terminal-Bench 3.0 moving from 4.6 to 28.3. Cybersecurity moved further than Z.ai says it expected, with CyberGym reaching 84.5%. Weights are not public yet.

    Is It Deployable?

    Partially, GLM-5.3 is live through the Z.ai API, the GLM Coding Plan, and ZCode. Weights are not out. Z.ai says it will publish them roughly two weeks after launch, once safety evaluation and hardening finish.

    • Which companies can move now: Startups and mid-market engineering orgs can adopt it today via the Coding Plan or API. Enterprises with data-residency or vendor-review rules should wait for weights. Security vendors and MSSPs get the most signal, and the most policy exposure.
    • Industries: Developer tooling, cloud infrastructure, application security, fintech and e-commerce engineering, and vendors shipping kernels, browser engines, or network stacks.
    • Applications: Repository-scale refactors, long-horizon CLI agents, CI failure triage, white-box vulnerability discovery, crash triage, and secure code review.

    Coding Results

    Terminal-Bench 3.0 moves from 4.6 to 28.3 against GLM-5.2. DeepSWE v1.1 moves from 46.2 to 66.9. Agents’ Last Exam (CLI) moves from 23.8 to 28.5. On GDPval-AA v2, which spans 44 occupations, GLM-5.3 scores 1,769.

    On Z.ai Code Bench, an internal evaluation, the company reports a 50% improvement over GLM-5.2. It reports 31.4% at roughly 50,000 output tokens per task. Claude Opus 4.8 scores 29.5% at 120,000 tokens. Claude Fable 5 still leads at 39.5% at maximum effort. Z.ai argues a private benchmark reduces contamination risk.

    aistudios

    On public suites, GLM-5.3 trails GPT-5.6 Sol and Fable 5 on several harder coding evaluations. All figures are vendor-reported, with harness, context length, and sampling settings documented in the announcement.

    The Cybersecurity Result

    Z.ai flags this one as unplanned. It added vulnerability-discovery data expecting better single-bug reasoning. Instead, capability kept compounding as training scaled. The model began forming coherent plans across complete exploitation chains.

    CyberGym, which tests discovery and validation from white-box source, moves from 77.2% to 84.5%. That edges past Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. ExploitBench, which requires root-cause reasoning and a working exploit, moves from 24.4% to 54.4%. Mythos 5 sits at 78.0%. On ExploitGym, GLM-5.3 completes 105 tasks in two hours and 130 in six. GLM-5.2 completes 29 and 39. Mythos 5 completes 181 and 247.

    The pattern is consistent. The deeper into the exploitation chain a benchmark sits, the larger the gain over GLM-5.2. The gap to closed frontier models also widens.

    Interactive Explainer

    Key Takeaways

    • GLM-5.3 reuses the GLM-5.2 base model; all gains come from post-training scaling.
    • Terminal-Bench 3.0 moves from 4.6 to 28.3; DeepSWE v1.1 from 46.2 to 66.9.
    • CyberGym hits 84.5%, ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%).
    • ExploitBench more than doubles to 54.4%, but trails Mythos 5 at 78.0%.
    • Weights ship in about two weeks, after safety evaluation and hardening.

    Check out the Z.ai GLM-5.3 technical blog, Zai_org announcement, Z.ai Security Disclosure Ledger and zai-org/GLM-5 on GitHub. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

    Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

    Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

    binance
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Fintech Fetch Editorial Team
    • Website

    Related Posts

    MG Ship adds AI route optimisation as logistics returns accelerate

    MG Ship adds AI route optimisation as logistics returns accelerate

    September 8, 2026
    IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B

    IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B

    September 7, 2026
    System helps humans predict when self-driving cars will make mistakes | MIT News

    System helps humans predict when self-driving cars will make mistakes | MIT News

    September 6, 2026
    M&T Bank expands enterprise AI after years of technology overhaul

    M&T Bank expands enterprise AI after years of technology overhaul

    September 5, 2026
    Add A Comment

    Comments are closed.

    Join our email newsletter and get news & updates into your inbox for free.


    Privacy Policy

    Thanks! We sent confirmation message to your inbox.

    notion
    Latest Posts
    Myriad: Who will top Spotify in 2026? Click to make your prediction.

    Your LG TV May Continue to Listen to You—Even When It Appears Turned Off

    September 8, 2026
    Cointelegraph

    Ethereum Outlines Key Focus Areas for the Upcoming Hegotá Upgrade

    September 8, 2026
    Why Jumia Technologies Stock Blasted Almost 16% Higher Last Month

    Why Jumia Technologies Shares Soared Nearly 16% Last Month

    September 8, 2026
    MG Ship adds AI route optimisation as logistics returns accelerate

    MG Ship adds AI route optimisation as logistics returns accelerate

    September 8, 2026
    AI Trading Bot Tried Day Trading (hint, it worked) AI Agent Tutorial

    AI Trading Bot Tried Day Trading (hint, it worked) AI Agent Tutorial

    September 8, 2026
    synthesia
    LEGAL INFORMATION
    • Privacy Policy
    • Terms Of Service
    • Social Media Disclaimer
    • DMCA Compliance
    • Anti-Spam Policy
    Top Insights
    Make Money With AI Is a Lie (Do This Instead)

    Make Money With AI Is a Lie (Do This Instead)

    September 9, 2026
    Cointelegraph

    Bitcoin Struggles to Maintain Support at $78,300 Amid Middle East Tensions Pressuring Risk Assets

    September 9, 2026
    bybit
    Facebook X (Twitter) Instagram Pinterest
    © 2026 FintechFetch.com - All rights reserved.

    Type above and press Enter to search. Press Esc to cancel.