An embeddable AI assistant bubble for any web app.
Self-hostable with Ollama. Open-source, no tracking, no data leaves your infrastructure.
lume-demo-doc.mp4
<LumeProvider
model="qwen2.5"
systemPrompt="You are a support assistant for Acme."
knowledgeBase={docs}
actions={actions}
components={components}
>
<App />
</LumeProvider>Drop <AssistantWidget /> anywhere inside. A floating bubble appears in the corner of your app. Click it — a chat panel slides up. The assistant already knows what page the user is on, can answer questions from your documentation, trigger real actions in your app, and render structured UI components inline in the conversation.
- Floating bubble UI — fixed-position overlay, zero impact on your app's layout
- Self-hostable — runs entirely on your machine via Ollama, no cloud required
LumeProvider— set up once at the root, control context from anywhere withuseLume()- Two-step intent classification — fast classifier routes each message before generating, dramatically more reliable than single-shot prompting
- Context injection — pass
{ page, plan, userId }and the assistant always knows where the user is - Knowledge base (RAG) — provide your docs as plain text, relevant chunks retrieved per query
- Actions — the assistant triggers real functions in your app, with a confirmation step before execution
- Structured UI components — the assistant renders your own React components inline instead of plain text
- Memory — persist conversation history across sessions with localStorage
- Proactive triggers — register conditions that open the widget and send a message automatically
- Suggestion chips — pre-defined prompts shown before the first message
- Connection status — widget shows live Ollama connection status in the header
- Debug mode — shows intent classification badge on each message for fast debugging
- Streaming responses — token-by-token output with a stop button
- Fully typed — complete TypeScript support with prop autocomplete
- Style isolated — inline styles only, no CSS leakage into or from the host app
- Zero dependencies — only React and Ollama required
- How it works
- Requirements
- Installation
- Quick start
- LumeProvider
- Configuration
- Debug mode
- Self-hosting
- How the AI adapts to your app
- useAssistant hook
- Contributing
- License
Lume doesn't train or fine-tune any model. Every message goes through two steps:
Step 1 — Intent classification (fast, ~200–400ms)
A lightweight Ollama call classifies the message into one of three categories before generating any response:
"take me to billing" → action:redirectTo
"what is my billing status" → component:billing_summary
"how do I reset my password" → text
This makes Lume dramatically more reliable than single-shot prompting. The classifier handles paraphrasing naturally — "am I close to my limit", "how much quota is left", and "show my usage" all correctly route to the same component.
Step 2 — Focused generation
Based on the classified intent, Lume builds a small, purpose-built prompt with one job:
text intent → persona + context + RAG — no JSON pressure
action intent → only the matched action schema — extract parameters
component intent → only the matched component schema — fill props from context
Each prompt is small and focused. The model never tries to be a classifier, JSON generator, and conversation partner at the same time.
The full prompt has six layers:
Layer 1 — Persona your systemPrompt
Layer 2 — Page context your context object
Layer 3 — RAG chunks most relevant docs from knowledgeBase
Layer 4 — Rules conciseness and formatting instructions
Layer 5 — Action schema only injected when intent is an action
Layer 6 — Component schema only injected when intent is a component
- Node.js 18+
- React 17+ (peer dependency)
- Ollama running locally
npm install @ovt2/lumeMinimal — chat only:
import { AssistantWidget } from '@ovt2/lume'
function App() {
return (
<>
{/* your app */}
<AssistantWidget
systemPrompt="You are a helpful assistant for this app."
/>
</>
)
}Recommended — with LumeProvider:
import { LumeProvider, useLume, AssistantWidget, defineAction, defineComponent } from '@ovt2/lume'
export default function App() {
const [page, setPage] = useState('dashboard')
const actions = [
defineAction(
'redirectTo',
{
description: 'Navigate to a page in the app',
parameters: {
page: { type: 'string', required: true, description: 'Page name' },
},
},
async ({ page }) => setPage(page as string)
),
]
const components = [
defineComponent(
'billing_summary',
{
description: 'Show billing info when the user asks about their plan or payments',
props: {
plan: { type: 'string', required: true, description: 'Plan name' },
status: { type: 'string', required: true, description: 'active or past_due' },
amount: { type: 'string', required: false, description: 'Amount due' },
},
},
(props) => <BillingCard plan={String(props.plan)} status={String(props.status)} />
),
]
return (
<LumeProvider
model="qwen2.5"
systemPrompt="You are a support assistant for Acme."
knowledgeBase={docs}
actions={actions}
components={components}
memoryKey="lume:acme"
accentColor="#6366f1"
title="Acme Support"
>
<InnerApp page={page} />
</LumeProvider>
)
}
function InnerApp({ page }: { page: string }) {
const { setContext } = useLume()
useEffect(() => {
setContext({ currentPage: page, userPlan: 'Pro' })
}, [page])
return (
<>
{/* your app */}
<AssistantWidget />
</>
)
}LumeProvider is the recommended way to integrate Lume. Set up your config once at the root — then control context and the widget from anywhere in your component tree using useLume().
import { LumeProvider } from '@ovt2/lume'
<LumeProvider
model="qwen2.5"
systemPrompt="You are a support assistant."
knowledgeBase={docs}
actions={actions}
components={components}
memoryKey="lume:myapp"
accentColor="#6366f1"
title="Support"
>
<App />
</LumeProvider>LumeProvider accepts all the same props as AssistantWidget. Props passed directly to AssistantWidget always win over provider values.
Access controls and context updates from anywhere inside the provider:
import { useLume } from '@ovt2/lume'
function MyPage() {
const { setContext, mergeContext, pushContext, open } = useLume()
// replace entire context on navigation
useEffect(() => {
setContext({ currentPage: 'billing', userPlan: user.plan, teamSize: team.length })
}, [page, user.plan, team.length])
// update a single key without touching others
const handleUpgrade = (newPlan: string) => {
mergeContext({ userPlan: newPlan })
pushContext(`User upgraded to ${newPlan}`)
}
// push live events — invisible to user, informs the assistant
const handleError = (err: Error) => {
pushContext(`Payment failed: ${err.message}`)
}
return <button onClick={() => open()}>Open assistant</button>
}| Method | Description |
|---|---|
setContext(ctx) |
Replace the entire context object |
mergeContext(partial) |
Update specific keys without replacing the whole context |
pushContext(note) |
Inject a hidden note into the conversation — never shown to the user |
open() |
Open the chat panel |
close() |
Close the chat panel |
toggle() |
Toggle open/closed |
clearHistory() |
Clear all messages |
All props are optional when using LumeProvider.
| Prop | Type | Default | Description |
|---|---|---|---|
model |
string |
'qwen2.5' |
Ollama model to use |
systemPrompt |
string |
'You are a helpful assistant.' |
Persona and instructions |
context |
object |
— | Current page/state — injected into every prompt |
knowledgeBase |
KnowledgeChunk[] |
[] |
Docs for RAG retrieval |
actions |
Action[] |
[] |
App functions the assistant can trigger |
components |
ComponentDefinition[] |
[] |
React components the assistant can render |
memoryKey |
string |
— | localStorage key for conversation persistence |
triggers |
Trigger[] |
[] |
Conditions that open the widget automatically |
suggestions |
string[] |
[] |
Prompt chips shown before the first message |
accentColor |
string |
'#6366f1' |
Bubble and send button color |
position |
'bottom-right' | 'bottom-left' | 'top-right' | 'top-left' |
'bottom-right' |
Corner to anchor to |
title |
string |
'Assistant' |
Label shown in the panel header |
debug |
boolean |
false |
Show intent classification badge on each message |
Pass any key/value object. It gets serialized into every prompt:
<AssistantWidget
context={{
currentPage: 'Invoice #1042',
userPlan: 'Pro',
userRole: 'admin',
lastError: 'export_failed',
}}
/>The assistant receives this as:
Current app context:
- current page: Invoice #1042
- user plan: Pro
- user role: admin
- last error: export_failed
With LumeProvider, use setContext or mergeContext from useLume() instead — no prop drilling needed.
Provide your app's documentation as an array of { title, content } chunks. On every message, the top 3 most relevant chunks are injected using keyword-frequency scoring — no embeddings, no vector database.
const docs = [
{
title: 'Exporting data',
content: `
You can export all your data as CSV from Settings → Export.
Exports are processed in the background and emailed when ready.
Large exports may take up to 10 minutes.
`,
},
]
<AssistantWidget knowledgeBase={docs} />Chunking tips:
- One topic per chunk — don't put your entire docs in one string
- Ideal chunk size: 150–400 words
- The title is also scored — make it descriptive
Actions let the assistant trigger real functions in your app. When the user's message matches an action, a confirmation card appears. The user confirms or cancels — then the handler fires.
Use defineAction to register an action:
import { defineAction } from '@ovt2/lume'
const actions = [
defineAction(
'redirectTo',
{
description: 'Navigate to a page in the app',
parameters: {
page: {
type: 'string',
required: true,
description: 'dashboard, billing, or settings',
},
},
},
async ({ page }) => router.push(page as string)
),
defineAction(
'inviteTeamMember',
{
description: 'Invite a new team member by email',
parameters: {
email: { type: 'string', required: true, description: 'Email address' },
role: { type: 'string', required: true, description: 'admin or viewer' },
},
},
async ({ email, role }) => {
await api.inviteUser({ email: email as string, role: role as string })
}
),
]Optional — add examples to improve intent matching:
defineAction(
'redirectTo',
{
description: 'Navigate to a page in the app',
examples: [
'take me to billing',
'go to settings',
'open the dashboard',
],
parameters: { ... },
},
async ({ page }) => router.push(page as string)
)How it works:
- User types
"invite alex@company.com as admin" - Classifier routes to
action:inviteTeamMember(~200ms) - Lume extracts parameters from the message
- Confirmation card appears
- User confirms → handler fires. User cancels → dismissed
Rules:
- Actions only trigger when the user explicitly asks — never on greetings or questions
- The confirmation step always happens — the model can never fire a handler silently
Instead of plain text, the assistant can render your own React components inline in the chat. Define the type, props schema, and a render function — Lume handles the routing and rendering.
Use defineComponent to register a component:
import { defineComponent } from '@ovt2/lume'
const components = [
defineComponent(
'billing_summary',
{
description: 'Show billing info when the user asks about their plan, payments, or subscription',
props: {
plan: { type: 'string', required: true, description: 'Plan name e.g. Pro' },
status: { type: 'string', required: true, description: 'active, cancelled, or past_due' },
nextBilling: { type: 'string', required: false, description: 'Next billing date' },
amount: { type: 'string', required: false, description: 'Amount due e.g. $49.00' },
},
},
(props) => (
<BillingSummaryCard
plan={String(props.plan)}
status={String(props.status)}
nextBilling={props.nextBilling ? String(props.nextBilling) : undefined}
/>
)
),
]How it works:
- User types
"what is my billing status?" - Classifier routes to
component:billing_summary(~200ms) - Lume fills props from app context
render(props)is called and the result appears as an assistant bubble- Component persists in conversation history — conversation continues normally
Key principle: Lume ships with zero built-in components. Every component is defined and rendered by your app. Lume routes the intent, fills the props, and calls render(). The visual output is 100% yours.
Enable conversation persistence so users return to their previous conversation after closing the widget or refreshing the page.
<LumeProvider
memoryKey="lume:myapp"
...
>Or on the widget directly:
<AssistantWidget memoryKey="lume:myapp" />How it works:
- Conversation history is saved to localStorage under the
memoryKey - On next load, the previous messages are restored automatically
- Clearing history via the trash icon also clears localStorage
- Only user and assistant messages are persisted — system messages are not stored
- History is capped at the last 50 messages to avoid localStorage bloat
Key naming: Use a unique key per app to avoid collisions if multiple Lume integrations share the same domain:
memoryKey="lume:acme-support" // good
memoryKey="lume" // too generic if multiple apps share domainMemory is opt-in — if memoryKey is not set, nothing is persisted and every session starts fresh.
Register conditions that automatically open the widget and send a message when something meaningful happens in your app — without the user having to ask.
<LumeProvider
triggers={[
{
id: 'billing-page',
condition: (ctx) => ctx.currentPage === 'billing',
message: "Need help understanding your invoice? I can walk you through it.",
delay: 2000,
},
{
id: 'payment-failed',
condition: (ctx) => ctx.lastError === 'payment_failed',
message: "It looks like your payment failed. Want me to help troubleshoot?",
},
{
id: 'high-usage',
condition: (ctx) => Number(ctx.usagePercent) >= 90,
message: "You're getting close to your plan limit. Want me to help you understand your options?",
},
]}
...
>Trigger shape:
interface Trigger {
id: string // unique — prevents double-firing
condition: (ctx: Record<string, string | number | boolean | null | undefined>) => boolean
message: string // what the assistant says
delay?: number // ms to wait before firing
}How it works:
- Developer registers triggers with conditions against the live context
- Every time
setContextormergeContextis called, all unfired triggers are evaluated - When a condition flips from
falsetotrue— the trigger fires - Lume opens the widget, waits for it to be visible, then inserts the message as an assistant bubble
- The trigger is marked as fired and never runs again for the rest of the session
Rules:
- Each trigger fires once per session — never repeats even if the condition stays true
- The message is pre-written by you — not AI generated. The AI takes over the moment the user replies
delayis the pause before the message appears after the condition is met — useful to avoid immediately interrupting the user on page load
Good trigger candidates:
// user lands on a page that commonly causes confusion
condition: (ctx) => ctx.currentPage === 'billing'
// something just went wrong
condition: (ctx) => ctx.lastError === 'export_failed'
// user is approaching a limit
condition: (ctx) => Number(ctx.usagePercent) >= 90
// first time a feature is used
condition: (ctx) => ctx.hasCompletedOnboarding === falseBad trigger candidates — avoid these:
// too broad — fires on almost every navigation
condition: (ctx) => ctx.currentPage !== 'dashboard'
// fires repeatedly as usage changes
condition: (ctx) => Number(ctx.usagePercent) > 50Show pre-defined prompt chips above the input bar before the user sends their first message. Tapping a chip sends it immediately.
<LumeProvider
suggestions={[
'What is my billing status?',
'Who is on my team?',
'How do I reset my password?',
'Show my usage',
]}
...
>Chips disappear after the first message is sent and reappear if the user clears the conversation history.
When not using LumeProvider, attach a ref for programmatic control:
import { useRef } from 'react'
import { AssistantWidget } from '@ovt2/lume'
import type { AssistantHandle } from '@ovt2/lume'
const ref = useRef<AssistantHandle>(null)
<AssistantWidget ref={ref} ... />| Method | Description |
|---|---|
ref.current.open() |
Open the chat panel |
ref.current.close() |
Close the chat panel |
ref.current.toggle() |
Toggle open/closed |
ref.current.pushContext(note) |
Inject a hidden context note into the conversation |
ref.current.clearHistory() |
Clear all messages and stop any active stream |
With LumeProvider, use useLume() instead — no ref needed.
Enable debug mode to see intent classification on every message:
<AssistantWidget debug />
// or at the provider level
<LumeProvider debug ...>Each user message shows a small badge:
- Purple
action:redirectTo— matched an action - Teal
component:usage_bar— matched a component - Gray
text— plain conversation
If something isn't triggering correctly, the badge tells you immediately whether the problem is in the classifier (wrong intent shown) or the generator (correct intent, wrong output). Remove debug before shipping to production.
Lume runs entirely on your own infrastructure. No data leaves your machine.
1. Install Ollama:
# macOS / Linux
curl -fsSL https://ollama.com/install.sh | sh
# Windows: download from https://ollama.com/download2. Pull a model:
ollama pull qwen2.53. Start Ollama:
ollama serve
# Listening on http://localhost:114344. Configure Lume:
<AssistantWidget model="qwen2.5" systemPrompt="You are a helpful assistant." />Lume connects to http://localhost:11434 by default. The widget header shows a live connection status dot — green when connected, yellow while checking, red when Ollama is unreachable with a tooltip showing run: ollama serve.
If your app is served from a different origin than localhost:
OLLAMA_ORIGINS="*" ollama serve
# or for a specific domain:
OLLAMA_ORIGINS="https://myapp.com" ollama serve| Model | Size | RAM | Best for |
|---|---|---|---|
qwen2.5 |
7B | 6 GB | Recommended — best tool-calling reliability for actions and components |
llama3.1 |
8B | 8 GB | Strong reasoning, good alternative to qwen2.5 |
mistral |
7B | 8 GB | Solid general-purpose option |
gemma3 |
4B | 4 GB | Lighter and faster — good for chat-only without actions or components |
Note: If you are using actions or structured UI components,
qwen2.5is strongly recommended. It has explicit tool-calling fine-tuning and produces reliable JSON output. Other models may work but are more prone to formatting errors.
Apple Silicon (M1/M2/M3): Ollama uses the GPU via Metal automatically. A MacBook Air M2 with 16GB runs qwen2.5 (7B) comfortably at ~40–60 tokens/sec.
The AI has no prior knowledge of your app. It adapts entirely through what you pass — no training, no fine-tuning:
systemPrompt → who the assistant is and how it should behave
context → what is happening right now in your app
knowledgeBase → how your app works (documentation)
actions → what the assistant can do
components → how the assistant can respond visually
memoryKey → conversation history across sessions
triggers → when the assistant speaks first
suggestions → prompt chips to guide first-time users
pushContext → live events as they happen (errors, state changes)
For building completely custom UIs — full control over rendering while keeping all core logic:
import { useAssistant } from '@ovt2/lume'
function MyCustomChat() {
const {
messages,
isStreaming,
error,
confirmation,
send,
stop,
clearHistory,
pushContext,
confirmAction,
cancelAction,
receiveProactive,
} = useAssistant({
model: 'qwen2.5',
systemPrompt: 'You are a helpful assistant.',
context: { page: 'Dashboard' },
knowledgeBase: docs,
actions: myActions,
components: myComponents,
memoryKey: 'lume:myapp',
})
return (
// your own UI here
)
}| Field | Type | Description |
|---|---|---|
messages |
Message[] |
Full conversation history including component messages |
isStreaming |
boolean |
True while a response is streaming |
error |
string | null |
Last error message, if any |
confirmation |
ConfirmationState |
Current action confirmation state |
send(text) |
function |
Send a user message |
stop() |
function |
Abort the current stream |
clearHistory() |
function |
Reset the conversation and clear localStorage |
pushContext(note) |
function |
Inject a hidden system note |
confirmAction() |
function |
Execute the pending action |
cancelAction() |
function |
Dismiss the pending action |
receiveProactive(msg) |
function |
Insert an assistant message without an Ollama call |
interface Message {
role: 'user' | 'assistant' | 'system'
content: string
intent?: string // debug only — e.g. "action:redirectTo", "component:usage_bar", "text"
componentCall?: {
component: string // matches a type in your components array
props: Record<string, unknown>
}
}type ConfirmationState =
| { status: 'idle' }
| { status: 'pending'; call: ActionCall; action: Action }
| { status: 'executing'; call: ActionCall; action: Action }Contributions are welcome. To run the project locally:
git clone https://github.com/ovt2/lume
cd lume
npm install
npm run devTo build the package:
npm run buildPlease open an issue before submitting a large pull request.
MIT — free to use, modify, and distribute.