TopheadlinesTopheadlines
    What's Hot

    Disgraced ex-NBC star Matt Lauer, fired ex-CNN boss Jeff Zucker and ‘killer’ Alec Baldwin all arrive for NYC wedding of ousted ex-CNN host Don Lemon

    April 6, 2024

    Over in a flash! Britain’s fun in the sun draws to a close as clear skies and 26C heat make way for thunderstorms and torrential rain

    May 12, 2024

    Chelsea 0-0 Preston – FA Cup third-round: Live score, team news and updates as Mauricio Pochettino’s side open up their campaign plus updates from the rest of Saturday’s late kick-offs with Aston Villa in action

    January 6, 2024
    Facebook Twitter Instagram
    Trending
    • Simple blood markers that could detect ALS years before diagnosis
    • WNBA’s furious trans debate deepens as Seattle star takes bitter swipe at Sophie Cunningham… and protestors gather to support Indiana Fever icon
    • Cookies sold at major grocery store recalled over fears of life-threatening reaction
    • UNLV basketball program ‘not for sale’: Runnin’ Rebels coach Josh Pastner walks back wild claim
    • When Suzie hit her 40s, she put her thinning hair, irritability, wired brain and palpitations down to the perimenopause. Then she started to mysteriously lose weight…
    • Oil tumbles below $90 after pause in fighting between the US and Iran but experts still fear rates hike
    • Tour de France 2026: Pogacar makes history, cycling steps up anti-doping, and popularity grows
    • Trendy meat diet sparks explosion of crippling medieval disease in young people
    Facebook Twitter Instagram
    TopheadlinesTopheadlines
    • Latest News

      Inside Epstein’s Zorro Ranch ‘where paedophile looked to carry out human experiments and create super-race breeding facility’ and ‘buried girls who were strangled during sex’

      February 24, 2026

      Cheating Utah ‘grief author’ heard sobbing down phone to 911 operator after allegedly murdering husband with poisoned Moscow mule

      February 23, 2026

      Victim reveals how she was left with flesh-eating disease after GP did not see her face-to-face then fled to India

      February 23, 2026

      ‘Andrew, the prince of darkness’: Global media mark ‘the end of privilege’ for Mountbatten-Windsor and gloat that the royal is ‘at rock bottom’

      February 22, 2026

      Melania Trump stuns in silver pants as she arrives at controversial Governor’s Dinner with husband Donald as dozens threaten boycott after president’s turbulent week

      February 22, 2026
    • Politics

      Law enforcement says eight killed by avalanche in California mountains | Weather News

      February 18, 2026

      Bangladesh PM-to-be and lawmakers sworn into parliament | Sheikh Hasina

      February 17, 2026

      Hillary Clinton gets in testy exchange with European leader over Trump

      February 16, 2026

      At least 11 Palestinians killed in Israeli attacks across Gaza | Gaza News

      February 15, 2026

      DOJ sends letter to Congress with list of people named in Epstein files, including Trump: Report

      February 15, 2026
    • Tech

      Apple’s AI Wearables Expected to Lean Heavily on Visual Intelligence

      February 23, 2026

      What to Know About At-Home STI Tests: Pros, Cons, and Recommendations (2026)

      February 23, 2026

      10 Clever Home Appliance Innovations You’ll See in 2026

      February 22, 2026

      Runlayer is now offering secure OpenClaw agentic capabilities for large enterprises

      February 22, 2026

      Tesla's "cheaper" Cybertruck arrives at $59,990, still far from the $40K promise

      February 21, 2026
    • Business

      Oil tumbles below $90 after pause in fighting between the US and Iran but experts still fear rates hike

      July 27, 2026

      Chancellor urged to rule out pensions raid to stop repeat of damaging speculation under Reeves and back ‘people who do the right thing’

      July 24, 2026

      Does the risk of ruinous care bills deter YOU from spending or gifting money to beat inheritance tax?

      July 23, 2026

      This is the one step Andy Burnham could take now to make us all richer and happier – and it wouldn’t cost him a penny: RACHEL RICKARD STRAUS

      July 21, 2026

      Why will it take half a year to register our property purchase on the Land Registry and should we be concerned?

      July 17, 2026
    • Sports

      WNBA’s furious trans debate deepens as Seattle star takes bitter swipe at Sophie Cunningham… and protestors gather to support Indiana Fever icon

      July 29, 2026

      UNLV basketball program ‘not for sale’: Runnin’ Rebels coach Josh Pastner walks back wild claim

      July 28, 2026

      Tour de France 2026: Pogacar makes history, cycling steps up anti-doping, and popularity grows

      July 27, 2026

      The McCanns cheer on Madeleine’s brother Sean as he swims in Commonwealth Games nearly two decades after sister disappeared

      July 25, 2026

      Rockies vs. Brewers MLB picks: Keep riding same parlay in lone matinee Friday

      July 24, 2026
    • Health

      Simple blood markers that could detect ALS years before diagnosis

      July 29, 2026

      Cookies sold at major grocery store recalled over fears of life-threatening reaction

      July 28, 2026

      When Suzie hit her 40s, she put her thinning hair, irritability, wired brain and palpitations down to the perimenopause. Then she started to mysteriously lose weight…

      July 27, 2026

      Trendy meat diet sparks explosion of crippling medieval disease in young people

      July 26, 2026

      Scientists find potential cause of embarrassing condition that causes excessive sweating in 15 million people… paving way for new treatments

      July 25, 2026
    • Science

      Ever wonder where our math symbols came from? Here are their stories

      February 19, 2026

      Some dog breeds carry a higher risk of breathing problems

      February 19, 2026

      New study links early smartphone ownership to health risks

      February 18, 2026

      The Story of Stories traces the arc of storytelling across human history

      February 17, 2026

      Listen to the crackle of ‘mini-lightning’ on Mars

      February 16, 2026
    • Entertainment

      Paul McCartney and Wings Exhibit Set at Rock & Roll Hall of Fame

      February 18, 2026

      Warner Bros Latest Hollywood Studio To Warn Seedance Over AI Infringement

      February 18, 2026

      Cardi B on Stefon Diggs Relationship Status

      February 17, 2026

      Jelly Roll Will Receive Country Radio’s Humanitarian Award

      February 17, 2026

      X Down For Thousands In U.S. And UK

      February 16, 2026
    TopheadlinesTopheadlines
    Home»Tech»Z.ai debuts open source GLM-4.6V, a native tool-calling vision model for multimodal reasoning
    Tech

    Z.ai debuts open source GLM-4.6V, a native tool-calling vision model for multimodal reasoning

    December 9, 20257 Mins Read
    Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Tumblr Email
    Z.ai debuts open source GLM-4.6V, a native tool-calling vision model for multimodal reasoning
    Share
    Facebook Twitter LinkedIn Pinterest Email



    Chinese AI startup Zhipu AI aka Z.ai has released its GLM-4.6V series, a new generation of open-source vision-language models (VLMs) optimized for multimodal reasoning, frontend automation, and high-efficiency deployment.

    The release includes two models in "large" and "small" sizes:

    1. GLM-4.6V (106B), a larger 106-billion parameter model aimed at cloud-scale inference

    2. GLM-4.6V-Flash (9B), a smaller model of only 9 billion parameters designed for low-latency, local applications

    Recall that generally speaking, models with more parameters — or internal settings governing their behavior, i.e. weights and biases — are more powerful, performant, and capable of performing at a higher general level across more varied tasks.

    However, smaller models can offer better efficiency for edge or real-time applications where latency and resource constraints are critical.

    The defining innovation in this series is the introduction of native function calling in a vision-language model—enabling direct use of tools such as search, cropping, or chart recognition with visual inputs.

    With a 128,000 token context length (equivalent to a 300-page novel's worth of text exchanged in a single input/output interaction with the user) and state-of-the-art (SoTA) results across more than 20 benchmarks, the GLM-4.6V series positions itself as a highly competitive alternative to both closed and open-source VLMs. It's available in the following formats:

    • API access via OpenAI-compatible interface

    • Try the demo on Zhipu’s web interface

    • Download weights from Hugging Face

    • Desktop assistant app available on Hugging Face Spaces

    Licensing and Enterprise Use

    GLM‑4.6V and GLM‑4.6V‑Flash are distributed under the MIT license, a permissive open-source license that allows free commercial and non-commercial use, modification, redistribution, and local deployment without obligation to open-source derivative works.

    This licensing model makes the series suitable for enterprise adoption, including scenarios that require full control over infrastructure, compliance with internal governance, or air-gapped environments.

    Model weights and documentation are publicly hosted on Hugging Face, with supporting code and tooling available on GitHub.

    The MIT license ensures maximum flexibility for integration into proprietary systems, including internal tools, production pipelines, and edge deployments.

    Architecture and Technical Capabilities

    The GLM-4.6V models follow a conventional encoder-decoder architecture with significant adaptations for multimodal input.

    Both models incorporate a Vision Transformer (ViT) encoder—based on AIMv2-Huge—and an MLP projector to align visual features with a large language model (LLM) decoder.

    Video inputs benefit from 3D convolutions and temporal compression, while spatial encoding is handled using 2D-RoPE and bicubic interpolation of absolute positional embeddings.

    A key technical feature is the system’s support for arbitrary image resolutions and aspect ratios, including wide panoramic inputs up to 200:1.

    In addition to static image and document parsing, GLM-4.6V can ingest temporal sequences of video frames with explicit timestamp tokens, enabling robust temporal reasoning.

    On the decoding side, the model supports token generation aligned with function-calling protocols, allowing for structured reasoning across text, image, and tool outputs. This is supported by extended tokenizer vocabulary and output formatting templates to ensure consistent API or agent compatibility.

    Native Multimodal Tool Use

    GLM-4.6V introduces native multimodal function calling, allowing visual assets—such as screenshots, images, and documents—to be passed directly as parameters to tools. This eliminates the need for intermediate text-only conversions, which have historically introduced information loss and complexity.

    The tool invocation mechanism works bi-directionally:

    • Input tools can be passed images or videos directly (e.g., document pages to crop or analyze).

    • Output tools such as chart renderers or web snapshot utilities return visual data, which GLM-4.6V integrates directly into the reasoning chain.

    In practice, this means GLM-4.6V can complete tasks such as:

    • Generating structured reports from mixed-format documents

    • Performing visual audit of candidate images

    • Automatically cropping figures from papers during generation

    • Conducting visual web search and answering multimodal queries

    High Performance Benchmarks Compared to Other Similar-Sized Models

    GLM-4.6V was evaluated across more than 20 public benchmarks covering general VQA, chart understanding, OCR, STEM reasoning, frontend replication, and multimodal agents.

    According to the benchmark chart released by Zhipu AI:

    • GLM-4.6V (106B) achieves SoTA or near-SoTA scores among open-source models of comparable size (106B) on MMBench, MathVista, MMLongBench, ChartQAPro, RefCOCO, TreeBench, and more.

    • GLM-4.6V-Flash (9B) outperforms other lightweight models (e.g., Qwen3-VL-8B, GLM-4.1V-9B) across almost all categories tested.

    • The 106B model’s 128K-token window allows it to outperform larger models like Step-3 (321B) and Qwen3-VL-235B on long-context document tasks, video summarization, and structured multimodal reasoning.

    Example scores from the leaderboard include:

    • MathVista: 88.2 (GLM-4.6V) vs. 84.6 (GLM-4.5V) vs. 81.4 (Qwen3-VL-8B)

    • WebVoyager: 81.0 vs. 68.4 (Qwen3-VL-8B)

    • Ref-L4-test: 88.9 vs. 89.5 (GLM-4.5V), but with better grounding fidelity at 87.7 (Flash) vs. 86.8

    Both models were evaluated using the vLLM inference backend and support SGLang for video-based tasks.

    Frontend Automation and Long-Context Workflows

    Zhipu AI emphasized GLM-4.6V’s ability to support frontend development workflows. The model can:

    • Replicate pixel-accurate HTML/CSS/JS from UI screenshots

    • Accept natural language editing commands to modify layouts

    • Identify and manipulate specific UI components visually

    This capability is integrated into an end-to-end visual programming interface, where the model iterates on layout, design intent, and output code using its native understanding of screen captures.

    In long-document scenarios, GLM-4.6V can process up to 128,000 tokens—enabling a single inference pass across:

    • 150 pages of text (input)

    • 200 slide decks

    • 1-hour videos

    Zhipu AI reported successful use of the model in financial analysis across multi-document corpora and in summarizing full-length sports broadcasts with timestamped event detection.

    Training and Reinforcement Learning

    The model was trained using multi-stage pre-training followed by supervised fine-tuning (SFT) and reinforcement learning (RL). Key innovations include:

    • Curriculum Sampling (RLCS): Dynamically adjusts the difficulty of training samples based on model progress

    • Multi-domain reward systems: Task-specific verifiers for STEM, chart reasoning, GUI agents, video QA, and spatial grounding

    • Function-aware training: Uses structured tags (e.g., <think>, <answer>, <|begin_of_box|>) to align reasoning and answer formatting

    The reinforcement learning pipeline emphasizes verifiable rewards (RLVR) over human feedback (RLHF) for scalability, and avoids KL/entropy losses to stabilize training across multimodal domains

    Pricing (API)

    Zhipu AI offers competitive pricing for the GLM-4.6V series, with both the flagship model and its lightweight variant positioned for high accessibility.

    • GLM-4.6V: $0.30 (input) / $0.90 (output) per 1M tokens

    • GLM-4.6V-Flash: Free

    Compared to major vision-capable and text-first LLMs, GLM-4.6V is among the most cost-efficient for multimodal reasoning at scale. Below is a comparative snapshot of pricing across providers:

    USD per 1M tokens — sorted lowest → highest total cost

    Model

    Input

    Output

    Total Cost

    Source

    Qwen 3 Turbo

    $0.05

    $0.20

    $0.25

    Alibaba Cloud

    ERNIE 4.5 Turbo

    $0.11

    $0.45

    $0.56

    Qianfan

    GLM‑4.6V

    $0.30

    $0.90

    $1.20

    Z.AI

    Grok 4.1 Fast (reasoning)

    $0.20

    $0.50

    $0.70

    xAI

    Grok 4.1 Fast (non-reasoning)

    $0.20

    $0.50

    $0.70

    xAI

    deepseek-chat (V3.2-Exp)

    $0.28

    $0.42

    $0.70

    DeepSeek

    deepseek-reasoner (V3.2-Exp)

    $0.28

    $0.42

    $0.70

    DeepSeek

    Qwen 3 Plus

    $0.40

    $1.20

    $1.60

    Alibaba Cloud

    ERNIE 5.0

    $0.85

    $3.40

    $4.25

    Qianfan

    Qwen-Max

    $1.60

    $6.40

    $8.00

    Alibaba Cloud

    GPT-5.1

    $1.25

    $10.00

    $11.25

    OpenAI

    Gemini 2.5 Pro (≤200K)

    $1.25

    $10.00

    $11.25

    Google

    Gemini 3 Pro (≤200K)

    $2.00

    $12.00

    $14.00

    Google

    Gemini 2.5 Pro (>200K)

    $2.50

    $15.00

    $17.50

    Google

    Grok 4 (0709)

    $3.00

    $15.00

    $18.00

    xAI

    Gemini 3 Pro (>200K)

    $4.00

    $18.00

    $22.00

    Google

    Claude Opus 4.1

    $15.00

    $75.00

    $90.00

    Anthropic

    Previous Releases: GLM‑4.5 Series and Enterprise Applications

    Prior to GLM‑4.6V, Z.ai released the GLM‑4.5 family in mid-2025, establishing the company as a serious contender in open-source LLM development.

    The flagship GLM‑4.5 and its smaller sibling GLM‑4.5‑Air both support reasoning, tool use, coding, and agentic behaviors, while offering strong performance across standard benchmarks.

    The models introduced dual reasoning modes (“thinking” and “non-thinking”) and could automatically generate complete PowerPoint presentations from a single prompt — a feature positioned for use in enterprise reporting, education, and internal comms workflows. Z.ai also extended the GLM‑4.5 series with additional variants such as GLM‑4.5‑X, AirX, and Flash, targeting ultra-fast inference and low-cost scenarios.

    Together, these features position the GLM‑4.5 series as a cost-effective, open, and production-ready alternative for enterprises needing autonomy over model deployment, lifecycle management, and integration pipel

    Ecosystem Implications

    The GLM-4.6V release represents a notable advance in open-source multimodal AI. While large vision-language models have proliferated over the past year, few offer:

    • Integrated visual tool usage

    • Structured multimodal generation

    • Agent-oriented memory and decision logic

    Zhipu AI’s emphasis on “closing the loop” from perception to action via native function calling marks a step toward agentic multimodal systems.

    The model’s architecture and training pipeline show a continued evolution of the GLM family, positioning it competitively alongside offerings like OpenAI’s GPT-4V and Google DeepMind’s Gemini-VL.

    Takeaway for Enterprise Leaders

    With GLM-4.6V, Zhipu AI introduces an open-source VLM capable of native visual tool use, long-context reasoning, and frontend automation. It sets new performance marks among models of similar size and provides a scalable platform for building agentic, multimodal AI systems.



    Source link

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email

    Related Posts

    Apple’s AI Wearables Expected to Lean Heavily on Visual Intelligence

    February 23, 2026

    What to Know About At-Home STI Tests: Pros, Cons, and Recommendations (2026)

    February 23, 2026

    10 Clever Home Appliance Innovations You’ll See in 2026

    February 22, 2026
    Top Posts

    Simple blood markers that could detect ALS years before diagnosis

    July 29, 2026

    Francis Ngannou claims he is Man United’s ‘favourite heavyweight boxer’ as he teases Tyson Fury

    July 31, 2023

    Turn Your Favorite Pet Photos Into a Pawfect Portrait for Just $20

    July 31, 2023

    Mauricio Diazgranados Is a Botanist in a Hurry

    July 31, 2023
    Don't Miss
    Tech

    Try These Tricks to Free Up More Screen Real Estate on a Mac

    Tech 2 Mins Read

    Justin PotAfter doing this the dock will disappear, allowing you to use that space for…

    After Fierce Lobbying, Treasury Sets Rules for Billions in Hydrogen Subsidies

    January 4, 2025

    The Heartbreaking Truth About Elvis and Priscilla Presley’s Love Story

    January 8, 2026

    This microphone picks up sounds by watching them

    November 19, 2025
    Stay In Touch
    • Facebook
    • Twitter
    • Instagram
    About Us
    About Us

    Delivering timely and accurate news updates, our website keeps you informed on the latest events, politics, entertainment, and more. Dive into captivating articles crafted by our team of expert writers, providing a comprehensive view of the world. Stay ahead with our trusted news source.

    Facebook Twitter Instagram
    Our Picks

    Breast cancer fears for almost 1,500 women at ‘very high risk’ of disease after NHS blunder meant they weren’t invited for annual check-ups

    March 5, 2024

    Revealed: The full list of countries where 12,500 migrants have entered Australia under citizenship blitz

    February 22, 2025

    John Bolton: ‘In a second Trump term, we’d almost certainly withdraw from NATO’

    August 4, 2023
    © 2026 Designed by TopHeadlineSpot.
    • Business
    • Latest News
    • Politics
    • Health
    • Entertainment
    • Sports
    • Science
    • Tech

    Type above and press Enter to search. Press Esc to cancel.