Legal LLM platform

Quoqo Large Language Model

Specialized legal LLMs with private RAG and commercial APIs

Discover jurisdiction-specific legal models, test them against private documents, then subscribe to OpenAI-compatible APIs for production workflows.

33planned legal model catalog
BGE-M3default embedding target
per orgprivate document corpus
OpenAIcompatible API surface
Source-grounded

Private legal RAG

Upload case bundles, research notes, contracts, or policy documents. Retrieval is scoped to each organization.

Discovery

Model marketplace

Browse Supreme Court, domain, and future jurisdiction models with use cases, latency notes, and capabilities.

Evaluation

Sandbox before purchase

Try prompts, inspect retrieved passages, compare models, and see usage before provisioning an API key.

Integration

Commercial API

Subscribe, create scoped API keys, and integrate chat completions, embeddings, and RAG into customer apps.

Discover QLLM models

Start with the live Supreme Court model, then expand into domain-specific legal models as they come online.

All jurisdictions Indian law Constitutional Contracts RAG-ready

Sandbox with private sources

Ask QLLM directly or turn on private RAG after uploading source documents. Retrieved passages appear alongside the answer.

Ask QLLM

Private RAG off
Top 6 chunks
QLLM
Select a model to start.

Pay for what you use

Try the console playground after signing in. The Developer plan is $49 a month and includes 5M billable tokens; after that, usage is $1.00 per 1M billable tokens, input and output alike. Billable tokens are tokens × the model's weight: 1 for 3B and 8B models, 10 for 70B and 20 for 70B 32k. Conveyancing (Maverick) is available on Enterprise.

Enterprise

Custom

For legal teams with private models and support.

+ SSO and org roles + dedicated Qdrant region + custom model routing + support and SLAs
Contact

Stripe Checkout

Subscription checkout uses Stripe Billing, with Customer Portal for plan changes and payment methods.

Metered usage

Prompt, completion, embedding, retrieval, storage, latency, and backend estimate are tracked per request.

Playground

Every signed-in workspace gets a fixed playground token allowance in the console. API keys and document search need a subscription.

API keys and usage

Generate scoped keys for integrations, monitor usage, and copy OpenAI-compatible examples.

Usage meter
Loading usage...
0tokens used
0retrievals
0 Bstored sources
curl https://api.qllm.quoqo.app/v1/chat/completions \
  -H "Authorization: Bearer qllm_live_..." \
  -H "Content-Type: application/json" \
  -d '{"model":"qllm-sc-in-8b","rag":true,"messages":[{"role":"user","content":"Summarize my uploaded document"}]}'

Sign in to QLLM

Notices

Built with Llama

Some QLLM models are built on Meta Llama models and are used under their licences.