All Posts

What LinkedIn Learned Building an AI Data Agent for 300+ Weekly Users (2024)

by Rick Radewagen4 min read

The "Fix with AI" button accounts for 80% of sessions. That tells you everything about where AI actually adds value in data workflows.


LinkedIn's data team faced a familiar problem: analysts spending their time helping colleagues find and query data instead of doing analysis. Their answer, SQL Bot, now serves over 300 weekly active users. The real story is what they learned about where AI genuinely helps.

After more than a year of development, LinkedIn published refreshingly honest numbers: 53% of responses score as correct or near-correct on internal benchmarks. Far from the 90%+ you see on academic benchmarks. Also far more useful than those benchmarks suggest.

The Architecture: Multi-Agent Systems Meet Knowledge Graphs

SQL Bot runs on LangChain and LangGraph, routing question types to specialized handlers. The real innovation is not the agent framework. It is the knowledge graph feeding it. The graph connects:

  • DataHub metadata: Table schemas, field descriptions, partition keys, top values for categorical fields
  • Query logs: Table and field popularity, common join patterns from successful queries
  • Domain knowledge: Business context collected from users through the UI
  • Example queries: Certified notebooks from DARWIN, LinkedIn's internal analytics platform

The graph refreshes weekly. Tables are the central nodes, linked to columns, documentation, historical queries, and product areas.

The Four-Stage Query Pipeline

  1. Retrieve context: Gather 20 candidate tables with embedding retrieval, filtered by popularity
  2. Rank context: An LLM narrows to 7 tables, then picks relevant columns in two tiers
  3. Write query: Generate SQL with documented assumptions
  4. Fix query: Validate syntax, check for hallucinations, send errors to a "Researcher Agent" with table search tools

The last stage proved crucial. The query fixer cut invalid tables and columns from 23% to 1%, and compilation success rose from 88% to 96%.

The Numbers That Matter

LinkedIn benchmarked 133 internal questions:

Configuration Table Recall Column Recall Score 4+ (Correct/Near-Correct)
Schemas only 45% 24% 9%
Full system 78% 56% 53%

The jump from 9% to 53% correct comes almost entirely from the knowledge graph. Not from better prompting. Not from smarter agents. Vercel's d0 proves the same point at the extreme: their semantic layer eliminated 80% of agent complexity.

User Satisfaction vs. Technical Accuracy

Here is where it gets interesting. Only 53% of responses score as technically correct. Yet 95% of users rate the query accuracy as "Passes" or above, and 40% rate it "Very Good" or "Excellent." Netflix's LORE points to the same principle: trust matters more than raw accuracy for adoption.

The disconnect? Users value the process, not just the output. SQL Bot helps them discover tables, understand schemas, and iterate toward a correct query, even when the first attempt misses.

The "Fix with AI" Revelation

The most-used feature is not query generation. It is the "Fix with AI" button that appears when a query fails. It accounts for 80% of sessions, and the team called it "easy to develop." The lesson: find the high-ROI pain point before building ambitious text-to-SQL. Users who know some SQL need help with the last mile, not the first.

Integration Multiplied Adoption by 5-10x

SQL Bot launched as a standalone chatbot. Adoption was modest. Then they integrated it into DARWIN, where analysts already work: a sidebar next to the query editor, the "Fix with AI" button on failures, persistent chat history, in-product feedback. Adoption grew 5 to 10 times.

The pattern repeats everywhere: AI tools succeed when embedded into existing workflows, not when they demand a context switch.

User Customization: Three Levers

LinkedIn built three ways for users to improve SQL Bot without involving the platform team:

  1. Dataset customization: Users map themselves to the datasets that matter to them
  2. Custom instructions: Free-form text that adds domain knowledge or guides behavior
  3. Example queries: Users index certified notebooks as few-shot examples

Self-serve customization is what lets one system cover many business verticals without a central team maintaining every domain.

How LinkedIn Compares to Uber's QueryGPT

Uber's QueryGPT tackles the same problem at similar scale, about 1.2 million interactive queries monthly: 300 daily active users, 78% of users say it saves time, and query authoring dropped from 10 minutes to 3.

Component LinkedIn SQL Bot Uber QueryGPT
Framework LangGraph + LangChain Multi-agent with specialized roles
Context Management Knowledge graph + DataHub Workspaces by business domain
Table Selection LLM re-ranking + clustering Intent Agent + Table Agent
Schema Handling Column relevance tiers Column Prune Agent
Validation Query fixer + Researcher Agent Execution + output validation

Neither system applies row-level security inside generated queries. For enterprises with strict data governance, that is a gap to check.

The Enterprise Benchmark Gap

Academic benchmarks like Spider show 90%+ accuracy. Spider 2.0, built to look like enterprise reality with 100-line queries and 1,000-column tables, shows the best models at 31%. Against that backdrop, LinkedIn's 53% on a data lake with millions of tables is strong. The lesson: benchmark on your own questions, not on Spider.

Practical Takeaways for Data Teams

1. Start with query debugging, not generation

80% of sessions, minimal development effort. Help users with the last mile first.

2. Invest in metadata before agents

The 9% to 53% jump came from the knowledge graph, not agent sophistication.

3. Integrate into existing workflows

Standalone chatbots stall. Embedded tools grow 5-10x.

4. Build for self-serve customization

Central teams cannot maintain every domain. Give users the levers.

5. Accept imperfect accuracy if the experience is valuable

53% technical accuracy with 95% satisfaction is not a contradiction. The journey delivers value even when the first answer misses.

6. Plan for validation and self-correction

Query fixers that catch hallucinations are table stakes. Budget for them from the start.

What's Next

LinkedIn's roadmap: faster responses, in-line query revisions, exposing the context SQL Bot used, and learning from user interactions. The most interesting thread is the shift from AI-generates-query to AI-helps-you-iterate. The future may look less like autonomous agents and more like copilots embedded where people already work.


SQL Bot shows that enterprise text-to-SQL is a systems engineering challenge, not a model capability problem. The winners invest in metadata, embed AI into existing workflows, and hit high-ROI pain points before chasing full automation.


If this excites you, we'd love to hear from you. Get in touch.

Sources:

Rick Radewagen

Rick is a co-founder of Dot, on a mission to make data accessible to everyone. When he's not building AI-powered analytics, you'll find him obsessing over well-arranged pixels and surprising himself by learning new languages.