//Intelligent AI Routing & Chat

SLM to LLM Switcher & Unified RAG Chat

If every user chat message triggers a $0.01 GPT-4o call, your SaaS loses its profit margin instantly. ButtrBase solves the AI unit economics crisis out of the box with Intelligent Model Routing.

1. The SLM to LLM Switcher

Instead of hardcoding gpt-4o or claude-3.5-sonnet, you request routing: "cost-optimized".

Our AI Gateway performs sub-millisecond heuristic complexity analysis: * SLM First (The Workhorse): Simple queries ("How do I reset my password?") are routed to blazingly fast, cheap Small Language Models like llama-3.1-8b or claude-3-haiku. * LLM Escalation (The Heavy Lifter): Complex reasoning or heavy coding queries are dynamically escalated to Large Language Models like gpt-4o.

This slashes your AI API bills by up to 80% with zero changes to your application logic.

2. The Unified search.chat() API

We combine our Hybrid Search (RAG) and our AI Gateway into a single API call. Instead of manually searching for context, formatting a prompt, and calling an LLM, you just do this:

const response = await client.search.chat("What is our refund policy?", {
  org_uuid: "org-123",
  routing: "cost-optimized" // Triggers the SLM -> LLM Switcher
});
ButtrBase automatically vector-searches the tenant's documents, formats the RAG prompt, picks the cheapest model capable of answering it, and returns the response.

3. Drop-in Component

To fully own the Developer Experience, we provide a drop-in Chat UI in our React, Solid, and Vue SDKs. You don't need to build chat bubbles, typing indicators, or streaming text parsers.

import { ButtrBaseChat } from '@buttrbase/react-sdk';

export function HelpBot() { return ( ); }