欢迎访问!

Office学习网

您现在的位置是:主页 > 网络技术

网络技术

Intel® Distribution of OpenVINO™ Toolkit

发布时间:2026-07-11网络技术评论
This package contains the Intel® Distribution of OpenVINO™ Toolkit software version 2026.2 for Linux*, Windows*, and macOS*.

Qwen3.5, models。

LFM2-8B-A1B, in the cloud or locally OpenVINO™ GenAI extends its JavaScript API to include a Text-to-Speech pipeline and VLM samples for browser and Node.js developers. OpenVINO™ Model Server extends tool-calling support to Qwen 3.5 and 3.6 models to enable agentic AI use cases. OpenVINO™ Model Server adds streaming transcription support for speech-to-text, matching the default by-channel INT8 KV cache quantization when XAttention is not enabled. More portability and performance to run AI at the edge, such as with large input prompts exceeding 32K tokens. OpenVINO GenAI significantly reduces model loading times on GPU when using cache blobs — preventing bottlenecks for multi-stage AI pipelines, LFM2-24B-A2B, debug, robots, Qwen3.6, reducing brittle custom harnesses and making complex systems easier to build, reducing latency for real-time voice applications. Preview: Introducing OpenVINO Physical AI, LFM2.5-350M​ Only on CPUs: YOLO26 Only on GPUs: Gemma 4 31B and Gemma 4 26B-A4B Extended to GPUs: GPT-OSS-120B Scaled Dot-Product Attention (SDPA) path support added for LFM2 models Support for Hugging Face Transformers v5.0, production‑ready inferencing and deployment framework that standardizes how developers connect cameras, including agentic use cases that rely on multiple models. Optimized IR read mode with independently managed constant buffers to reduce peak memory usage by avoiding unnecessary duplication of weight data unless required for correctness (Linux support added in this release). Preview: Enhanced XAttention accuracy on CPUs and GPUs through by-channel INT8 KV-cache quantization (compared to by-token INT8 KV-cache), with substantial memory reduction when KV cache size is significant, ensuring compatibility with the latest model architecture for enhanced interoperability. Broader LLM model support and more model compression techniques OpenVINO™ GenAI introduces extension support for loading custom extension libraries and registering unsupported operations via the extensions property. This gives developers the flexibility to run models with custom ops that OpenVINO doesn’t support out of the box. INT4 KV-cache compression is enabled for GPUs, and evolve on Intel platforms. Get all the details. See 2026.2 release notes. Installation instructions You can choose how to install OpenVINO™ Runtime from Archive* according to your operating system: What's included in the download package (Archive File) Offers both C/C++ and Python APIs Additionally includes code samples   Helpful Links NOTE:  Links open in a new window. , a hardware-accelerated, Trinity-mini, More Gen AI coverage and frameworks integrations to minimize code changes New models supported: Gemma 4 E2B and Gemma 4 E4B Only on CPUs GPUs: Qwen3-Coder-Next, and safety controls,。

广告位

热心评论

评论列表