Alibaba Cloud OSS: From Object Storage to AI-Native Data Infrastructure with Vector Bucket & Metaquery

简介: This article provides an in-depth look at how OSS builds an AI-native storage foundation to empower key scenarios including Retrieval-Augmented Generation (RAG), enterprise search, and AI-powered content management—helping you efficiently build next-generation intelligent applications.

By Justin See


As data volumes explode and AI becomes central to enterprise competitiveness, traditional object storage must evolve. Alibaba Cloud Object Storage Service (OSS) is leading this transformation—moving from a passive data carrier to an intelligent, AI-powered knowledge provider.

This blog explores OSS’s technical foundation, its evolution toward AI-native workloads, and deep dives into two major innovations: Vector Bucket and Meta Query.

In the AI era, Vector Data and the ability to query it, is the foundation of AI applications. This diagram illustrates how diverse unstructured data types—documents, images, videos —are transformed into dimensional vector representations through embedding. These vectors serve as the backbone for advanced AI capabilities such as content awareness and semantic retrieval. By converting raw data into structured vector formats, OSS enables intelligent indexing, search, and discovery across massive datasets, forming the basis for applications like retrieval-augmented generation (RAG), enterprise search, and AI content management.

image.png

1. OSS Architecture and Core Capabilities


Alibaba Cloud OSS is a distributed, highly durable object storage system designed for massive-scale unstructured data. It supports workloads ranging from web hosting and backup to AI training and semantic retrieval.

OSS Architecture Layers

Layer Components
Application Interfaces API, SDKs, CLI
OSS Service Control Layer, Access Layer, Metadata Indexing
Storage Infrastructure Erasure Coding, Zone-Redundant Storage
AI Integration MCP Server, AI Assistant, Semantic Retrieval Engine

Storage Classes and Cost Optimization

Class Use Case Retrieval Time
Standard Hot data Instant
IA Infrequent access Instant
Archive Cold data Minutes
Cold Archive Deep archive Hours
Deep Cold Archive Long-term retention Hours

Lifecycle policies and OSS Inventory now support predictive tiering and DryRun simulations to prevent accidental deletions.

Performance and QoS

OSS Resource Pool QoS enables:

● Shared throughput across buckets

● Priority-based scheduling (P1–P4)

● Guaranteed minimum bandwidth

● Real-time monitoring and dynamic allocation

This ensures stable performance across mixed workloads—online services, batch jobs, AI training, and data migration.

2. OSS Vector Bucket: AI-Native Storage for Embeddings


Vector data is the foundation of semantic search, recommendation engines, and retrieval-augmented generation (RAG). OSS Vector Bucket introduces native support for storing, indexing, and querying high-dimensional vectors.

image.png

Architecture

● Raw data stored in OSS Bucket

● Embedding service generates vectorized data

● Vectors stored in OSS Vector Bucket

● MCP Server enables semantic retrieval for AI Content Awareness

● AI Agent integrates with RAG and AI Semantic Retrieval

Difference with Traditional Vector Database

As enterprise AI workloads scale, the volume of vectorized data grows exponentially—driving up infrastructure costs and straining traditional storage architectures. OSS Vector Bucket offers a cost-effective alternative by decoupling compute from storage, allowing vector queries to be executed directly on OSS without relying on tightly coupled service nodes. In most AI use cases, customers are able to tolerate higher retrieval latencies—hundreds of milliseconds; OSS vector bucket has higher latencies but significantly reducing operational costs through pay-as-you-go pricing for both storage and query scans. By migrating to OSS Vector Bucket, enterprises can build retrieval-augmented generation (RAG) applications that are not only scalable and performant, but also financially sustainable.

Use Cases

● RAG Applications

● AI Agent Retrieval

● AI-powered Content Management Platform for Social Media, E-commerce, Media

Key Features

● Supports hundreds of billions of vectors per account

● Native OSS API with SDK, CLI, and console access

● Integrated with Tablestore for high-performance workloads

● Pay-as-you-go pricing: storage capacity, query volume

● Unified permission management via OSS bucket policies

Key Benefits

● Lower vector database costs

● Automatic scaling & limitless elasticity

● Reduces data silo and management overheads

Watch Demo: https://www.youtube.com/watch?v=xgY7k8hVS20

Documentation: Vector Bucket Documentation

3. OSS Metaquery: Content Awareness & Semantic Search


OSS Metaquery transforms raw unstructured data into an intelligent, searchable knowledge layer by automatically generating embeddings for every newly added object, eliminating the need for manual preprocessing or external vector pipelines. Once embedded, the data becomes immediately accessible through semantic search, allowing users to query their OSS buckets using natural language rather than rigid keywords or file paths. The system combines scalar filtering—such as metadata conditions on size, time, or tags—with vector-based similarity search to deliver highly relevant results through a hybrid retrieval engine. In production deployments, this approach has demonstrated precision‑recall rates of up to 85 percent, significantly outperforming traditional self‑built search solutions and enabling enterprises to build powerful retrieval‑augmented and content‑aware applications directly on top of OSS.

image.png

Use cases

● Intelligent Enterprise Search

● RAG applications

● AI-powered Content Management Platform

Key Features

● One-click enables automatic embeddings

● Semantic search using natural language

● No extra management overheads

Key Benefits

● Time to market significantly shortened without the need to build vector embedding & natural language search engine

● Higher accuracy search with multi-channel recall

● Lower costs by leveraging cost effective object storage, and no additional components to build or manage

Watch Demo: https://www.youtube.com/watch?v=xgY7k8hVS20

Documentation: Metaquery Documentation

4. OSS Accelerator and Other Tooling Upgrades


To support high-throughput AI workloads, OSS also introduced:

OSS Accelerator (EFC Cache) Upgrade

The OSS Accelerator introduces a high‑performance, compute‑proximate caching layer designed to dramatically improve data access speeds for AI, analytics, and real‑time workloads running on Alibaba Cloud. By deploying NVMe‑based cache nodes in the same zone as compute resources, the accelerator reduces read latency to single‑digit milliseconds and supports burst throughput of up to 100 GB/s with more than 100,000 QPS. It requires no application changes: all OSS clients, SDKs, and tools automatically benefit from acceleration through a unified namespace, with strong consistency maintained between cached data and the underlying OSS objects. Multiple caching strategies—including on‑read prefetch, manual prefetch, and synchronized prefetch—ensure that hot data is always available at high speed, while an LRU eviction policy manages cache capacity efficiently. This upgrade is particularly impactful for workloads such as model loading, BI hot table queries, and high‑frequency inference, enabling organizations to achieve low‑latency performance while keeping the majority of their data stored cost‑effectively across OSS Standard, IA, Archive, and Cold Archive tiers.

OSSFS V2 Upgrade

OSSFS V2 is a next‑generation, high‑performance mount tool designed to make OSS behave like a local file system for AI, analytics, and containerized workloads. Built with a lightweight protocol and deeply optimized I/O path, OSSFS V2 delivers substantial performance improvements over previous versions—achieving up to 3 GB/s single‑thread read throughput and significantly reducing CPU and memory overhead. It introduces a negative metadata cache to minimize redundant lookups, improving responsiveness for workloads that perform frequent directory scans or small‑file access. OSSFS V2 is fully compatible with Kubernetes environments through CSI integration, enabling seamless mounting of OSS buckets as persistent volumes in ACK and ACS clusters. This makes it ideal for applications that cannot be easily modified to use the OSS SDK directly, such as legacy data processing pipelines, distributed training frameworks, and containerized AI workloads. With strong consistency, elastic scalability, and support for high‑concurrency access, OSSFS V2 allows organizations to use OSS as a high‑throughput data layer across the entire AI lifecycle—from data ingestion and preprocessing to model training and inference.

OSS Connector for Hadoop V2 Upgrade

The OSS Connector for Hadoop V2 delivers a major performance and efficiency leap for data lake and big data analytics workloads running on Hadoop and Spark. This new version introduces an adaptive prefetching mechanism that eliminates redundant metadata operations, significantly improving read throughput—up to 5.8× faster in benchmark tests—and reducing end‑to‑end SQL query time by 28.5 percent. It also integrates seamlessly with OSS Accelerator, enabling hot data to be cached on NVMe storage close to compute nodes, which can further reduce query latency by up to 40 percent. Built on top of the next‑generation OSS Java SDK V2, the connector adopts default V4 authentication for stronger security and improved performance. These enhancements make OSS Connector for Hadoop V2 a high‑performance, cloud‑native storage interface for AI data lakes, supporting large‑scale ETL, interactive analytics, and machine learning pipelines with significantly lower overhead and higher throughput.

OSS Resource Pool QoS Upgrade

The OSS Resource Pool QoS upgrade introduces a unified performance management framework that allows multiple buckets and workloads to share a common throughput pool while maintaining predictable service quality. Instead of each bucket operating in isolation, enterprises can now allocate and prioritize throughput across business units, job types, or RAM accounts. QoS policies support both priority‑based dynamic control—ensuring critical online services receive the bandwidth they need during peak hours—and minimum guaranteed throughput, which protects lower‑priority batch or analytics jobs from starvation. This fine‑grained control enables stable performance for mixed workloads such as AI training, data preprocessing, online inference, and large‑scale data migration, all while maximizing overall resource utilization. With throughput pools scaling to tens of terabits per second, OSS QoS becomes a foundational capability for storage‑compute separation architectures and unified AI data lakes.

OSS SDK V2 Upgrade

The OSS SDK V2 upgrade delivers a comprehensive modernization of the OSS client experience, offering higher performance, stronger security, and broader language coverage for developers building AI, analytics, and cloud‑native applications. This new version introduces a fully asynchronous API architecture that significantly improves throughput for high‑concurrency workloads such as model training, data ingestion, and large‑scale ETL. It adopts default V4 authentication for enhanced security and more efficient request signing, reducing overhead for frequent or parallel operations. SDK V2 provides full language support—including Go, Python, PHP, .NET, Swift, and Java—ensuring consistent behavior and performance across diverse development environments. With improved error handling, streamlined configuration, and optimized network usage, SDK V2 enables developers to interact with OSS more efficiently while taking full advantage of the platform’s evolving AI‑native capabilities.

These upgrades enable OSS to serve as the backbone for AI data lakes and inference platforms.

Conclusion

Alibaba Cloud OSS is no longer just object storage—it’s a full-stack, AI-native data infrastructure. With Vector Bucket, Metaquery, Content Awareness, Semantic Retrieval, OSS Accelerator and other upgrades, OSS supports:

● AI training and inference

● Intelligent data discovery

● Semantic search and retrieval

● Cost-effective RAG applications

● AI-powered Content Management Platform

● Unified data lake architectures

Whether you're building AI Agents, managing enterprise digital assets, or scaling AI workloads, OSS is ready to power your next-generation data strategy.

相关实践学习
对象存储OSS快速上手——如何使用ossbrowser
本实验是对象存储OSS入门级实验。通过本实验,用户可学会如何用对象OSS的插件,进行简单的数据存、查、删等操作。
相关文章
|
9月前
|
人工智能 自然语言处理 Serverless
新突破!阿里云携手技威时代共同开启 IPC 智能化新阶段
阿里云IPC AI方案融合千问大模型视觉理解与OSS Metaquery多模态检索,实现Serverless、低成本、高准召的智能视频检索。无需硬件改造,存量设备即可升级,一句自然语言唤醒沉睡视频,让“看”升级为“懂”。
472 5
|
存储 人工智能 自动驾驶
高性能存储CPFS在AIGC场景的具体应用
高性能存储CPFS在AIGC场景的具体应用
|
3月前
Tushare接口文档:指数基本信息(index_basic)
本文旨在对Tushare的指数基本信息`index_basic`数据接口进行介绍,提供更多参考示例和使用说明。 通过本接口可获取指数的基础信息,例如指数代码、全称简称、发布方、基期、加权方式等。在指数行情等多个需要依赖`ts_code`而不清楚后缀的情况下,可以使用该接口获取指数代码`symbol`与`ts_code`的对应关系。本文还介绍了如何通过传递`market`和`category`字段分别指定发布指数的交易所或服务商以及指数分类。
357 1
|
5月前
|
存储 人工智能 自然语言处理
知识库接入还能这么玩?Tablestore 四种方式实战揭秘
本文详解 Tablestore 知识库服务 API 设计、四种接入方式、多维度评测结果及 PDS、ECS 等客户落地案例,助力企业快速集成高质量 RAG 能力。
1038 125
|
6月前
|
存储 人工智能 弹性计算
揭秘千问 APP 千万级 AI 订单背后的记忆存储实践
2026年春节,千问 APP “春节请客计划” 9 小时破 1000 万单,依赖 Tablestore 构建的一站式记忆系统:支持短期/长期记忆统一管理、毫秒级读写、Serverless 弹性伸缩、多模态数据融合及原生向量检索,实现数十亿条记忆的高效存储与实时流转。
1434 118
|
7月前
|
存储 人工智能 弹性计算
阿里云网盘 Skill 上线,附 OpenClaw 配置网盘空间实操教程
阿里云网盘正式上线OpenClaw专属Skill,为龙虾AI提供云端存储、多端实时同步与精细权限管控,解决本地空间不足、跨端难协同、数据不安全等痛点,3分钟配置即享高性价比(200GB/月仅6.6元)AI工作流升级。
1885 6
|
4月前
|
存储 人工智能 运维
3 人团队零推广获 1.2 万用户:Matrees 如何用 OSS 向量 Bucket 低成本构建 AI 创作平台
Matrees 是 Z 世代创作者的 AI 虚拟世界平台,3 人团队几乎零推广获取 1.2 万用户。依托阿里云 OSS Vector Bucket 实现全托管向量检索,成本降 90%,让创作者专注构建虚拟世界。
489 1
|
4月前
|
存储 数据采集 机器人
具身智能爆发背后,存储如何成为关键基础设施?
具身智能爆发式发展带来存储新挑战:多模态数据采集、仿真生成、清洗标注、分布式训练与低延迟推理各环节诉求迥异。阿里云以 OSS 统一底座(支持 QoS 管控、冷热分层、Vector Bucket 向量检索)+ CPFS 高性能训练引擎协同,构建全链路、高弹性、低成本存储方案。
764 0
具身智能爆发背后,存储如何成为关键基础设施?
|
4月前
|
存储 运维 数据管理
告别“大海捞针”:OSS Vector Bucket 如何赋能媒资管理平台
在 AI 时代,媒资平台面临多模态数据爆炸式增长的管理挑战。阿里云 OSS Vector Bucket 提供统一向量存储与语义检索能力,支持 30 亿级素材秒级精准查找,打破数据孤岛,降低成本,助力内容创作提效降本。
418 11
|
4月前
|
人工智能 自然语言处理 算法
800 家门店的便利零售如何低成本接入 AI 推荐?美好超市基于 OSS 向量 Bucket 的实践
本文讲述江苏头部连锁超市“美好超市”如何借助阿里云 OSS 向量 Bucket,零算法团队、免自建系统,快速实现语义搜索与智能推荐:购物车实时推荐、详情页智能搭配、自然语言问答购等场景全面升级,让 AI 推荐真正普惠中小企业。
385 0