LanceDB’s cover photo
LanceDB

LanceDB

Information Services

San Francisco, California 12,388 followers

The multimodal lakehouse for AI, accelerating large-scale data curation and feature engineering.

About us

LanceDB is the multimodal lakehouse for AI: the unified foundation for AI data, where every type lives together in open format in your own cloud. It unlocks AI curation, feature engineering, and training at scale on datasets too large and too mixed for a stitched-together stack to handle. Frontier labs and generative-media teams run their training, search, and curation on LanceDB, including Netflix, Runway, Midjourney, and World Labs.

Website
http://lancedb.com
Industry
Information Services
Company size
11-50 employees
Headquarters
San Francisco, California
Type
Privately Held
Founded
2022

Locations

Employees at LanceDB

Updates

  • Vector search that works at 10 million vectors shouldn’t need to be redesigned when the table grows to 10 billion. In our latest benchmarks, LanceDB reached 96.2% recall at 1,614 QPS on 10M vectors. At 10B vectors, it sustained more than 1,000 QPS with 18.05 ms p50 and 21.61 ms p99 latency. But growing search is more than query performance. The index has to keep up too. Yang Cen breaks down how LanceDB combines distributed and incremental indexing, query-time precision controls, and tiered caching to support search from millions to billions of vectors. Read the full post: https://lnkd.in/gXiakehz

    • No alternative text description for this image
  • A great blog written by Ayush Chaurasia & Aritra Roy Gosthipaty 🤗👏 https://lnkd.in/dMYGKCrT funes, by Hugging Face, turns past agent sessions into memory your agents can actually use. It indexes Claude Code, Codex, pi, and Hermes traces into one local Lance dataset, then gives the agent 'recall' and 'get' tools. The next time a task depends on old reasoning, the agent can pull the original passage back. No LLM summarizing your traces at ingest. Just your working record, local by default, shareable when you choose.

  • Feature of the Week: Index prewarm now reads in parallel byte windows We’re launching a new series, Feature of the Week, to spotlight what our engineers are building and the work behind it. This week’s feature was built by Yang Cen: https://lnkd.in/gXpaQdeE Instead of indexing one partition at a time and decoding on the same thread, now in Lance you can grab partitions in 64MB chunks and hand the decoding off to a separate CPU Pool. Loading a billion-vector index into memory used to take 95 minutes, now it takes 3. See you in next week’s video 👋

  • What if pretraining didn’t require turning your dataset into a chain of new datasets? Ayush Chaurasia trained GPT-2 from raw text in 25 minutes on 8 H100s using LanceDB. The pipeline goes from raw text → curation → tokenization → training, all in one LanceDB table. A few results: → 2.43B tokens prepared in 11 minutes → 124M GPT-2 trained in 14 minutes at 3.18M tokens/sec → Data streamed and globally shuffled directly from S3 at 3.16M tokens/sec, matching local disk → No pre-shuffled dataset Change the tokenizer, curation filter, or sequence length without rebuilding a chain of intermediate datasets. Read the full technical walkthrough: https://lnkd.in/g-2FAdkZ

    • No alternative text description for this image
  • Your coding agent forgets everything the moment a session ends. Ask it what it decided last week and why, and it has no idea. Funes, built by Hugging Face, indexes past sessions from Claude Code, Codex, pi, and Hermes into one shared memory your agent can query. → Local embeddings, vector + BM25 hybrid search, reranking, recency weighting → The memory itself is a Lance dataset → Same file whether you're querying locally or publishing it to the Hub for your team That memory travels with the dataset, not the session. 🔗 Check it out : https://lnkd.in/gkeeZVvz

  • Lance is now natively supported in 🤗 LeRobot. Go from ingestion and curation through training and eval on the same robotics dataset, without exporting or rematerializing between steps. With Lance-backed LeRobot datasets, you can: • Train directly from S3, GCS, or Hugging Face Storage Buckets • Keep using the same LeRobotDataset API • Globally shuffle training data while reading remotely • Avoid downloading the full dataset before training Same LeRobot training workflow. A storage format built for large multimodal datasets underneath it. Huge thanks to the Hugging Face LeRobot team Caroline Pascal and Quentin Lhoest for working with us to get this merged. https://lnkd.in/gnc7ZN3A

  • Physical AI systems are generating more data than ever. The harder problem is finding the tiny fraction of fleet experience that can actually make the model better. As models improve, the useful examples move further into the long tail. A rare grasp failure. An unusual pedestrian interaction. A scene visually similar to a failure nobody thought to label. That changes what the data infrastructure needs to do. In his latest post, LanceDB CTO Lei Xu explains why data mining is becoming one of the most important infrastructure problems in robotics and autonomous systems, and why feature engineering, retrieval, training, and evaluation increasingly need to operate on the same underlying multimodal data. Read the full post ↓ https://lnkd.in/gtDx6cVW

    • No alternative text description for this image
  • 🌟 Reverie, LanceDB's summit for AI builders, lands Nov 5 in San Francisco with speakers from NVIDIA, Runway, and Luma. → Apply to attend: https://lnkd.in/gzGaaFsk OSS Updates: 🔍 ACORN-1 cuts worst-case search latency by up to 250x in Lance 🌳 Lance segmented indexes now support bloom filters 🪞 LanceDB now supports materialized views 🐍 LanceDB adds a Function framework for scalar UDFs LanceDB Enterprise: ⚡ WAL compaction fetches SSTables 4x faster 🧹 Per-job scratch storage cuts cleanup time 58% 🔀 SQL now supports CREATE TABLE ... CLONE 🔗 More in August newsletter: https://lnkd.in/gA92STSn

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • View organization page for LanceDB

    12,388 followers

    Excited to welcome Fei Xue, Jeff Lane, and Cole Dumanski to the team! 🎉 Fei Xue is joining us as a Product Manager! Her career has been largely rooted in cloud, data, and AI. First starting out in Engineering, then moving into GTM, and now has spent the last 7+ years as a Product Manager – "A path that gives me both a builder's instinct for how things work and a customer-facing sense of why they matter." Outside of work, she loves to stay active, whether it's swimming laps, hitting the trails, or simply being outdoors. Jeff Lane is joining us to lead Revenue Operations! He brings 30+ years of experience in GTM, starting in Sales Engineering, Professional Services, IT and now in Revenue Operations. He's run GTM Ops end-to-end at early stage companies, built global centers of excellence for the GTM tech stack at scale, and supported field organizations of 1,000+ from demand gen to customer support. He recently moved to Bellingham, WA so you can find him taking full advantage of the outdoors! Cole Dumanski is joining us as an intern on the Marketing team! He's currently studying Management Sciences Engineering at University of Waterloo; previously interned at Theory Ventures and Upfront Ventures, working across data engineering, AI engineering, content and venture investing. Outside of work, you can find him playing hockey, practicing golf, doing random outdoor activities or watching Star Wars! Join the team → lancedb.com/careers

    • No alternative text description for this image

Similar pages

Browse jobs