QuickLLM - Guide

QuickLLM User Guide

QuickLLM keeps inference, conversations, documents, embeddings, retrieval, and generated text on your device while making model fit and network purpose visible.

Getting Started

  1. Run device diagnostics and choose a model that fits.
  2. Download, verify, and load the model.
  3. Start a conversation and choose a Profile.
  4. Optionally create a Knowledge Base, add trusted documents, and verify citations.
  5. Export important work or create an encrypted backup before moving or deleting data.
Tip: Choose a smaller context or model before forcing a configuration that does not fit memory.

Interface Overview

AreaPurpose
ChatConversation list, model/Profile/Knowledge selection, messages, citations, context, Send, and Stop.
ModelsCatalog, search, compatibility, downloads, installed files, Profiles, adapters, and runtime state.
KnowledgeKnowledge Bases, documents, extraction, indexing, retrieval tests, grounding, and citations.
Templates & ActionsReusable prompts, typed workflows, mobile share input, result review, and save paths.
UtilitiesDiagnostics, downloads, desktop API, backup/restore, quick action, Settings, Help, privacy, and licenses.

Models That Fit, Answers You Can Trace

QuickLLM distinguishes measured device capability from estimates and keeps grounded answers linked to local document excerpts instead of inventing unsupported citations.

  • Private local chat: Run conversations through packaged local inference with visible model, Profile, Knowledge, context, generation, and Stop state.
  • Model fit: Measure device capability, compare memory and storage needs, download verified models, and explain unsupported configurations.
  • Knowledge Bases: Ingest supported documents, build local indexes, retrieve grounded excerpts, and open citations at their source.
  • Profiles and LoRA: Save model, context, sampling, backend, and adapter combinations for repeatable work.
  • Templates and Actions: Reuse private prompts and run shared-text workflows that keep results available for copy, share, or conversation.
  • Portable control: Export conversations, run a desktop local API when explicitly enabled, and create encrypted backup and restore files.
Note: Downloaded models and source documents may have separate licenses and usage restrictions. Review them before use or redistribution.

Settings

Open Settings to review the following app-specific controls:

  • Native engine, backend, threads, context, cache, storage, and model preferences
  • Knowledge extraction, embeddings, retrieval, grounding, and citation behavior
  • Desktop local API, network, mobile share, templates, and actions
  • Privacy, encrypted backup, deletion, appearance, shortcuts, diagnostics, and licenses

Trial, Free, and Pro use the same functional feature set. Pro removes mobile ads and the daily desktop support reminder.

Keyboard Shortcuts

On Apple platforms, use Command where the app maps a displayed Ctrl shortcut to the platform primary modifier. The in-app shortcut page remains authoritative.

ShortcutAction
Cmd/Ctrl + NNew conversation
Cmd/Ctrl + KSearch conversations
Cmd/Ctrl + MChoose model
Cmd/Ctrl + BChoose Knowledge Base
Cmd/Ctrl + Alt/Option + IConversation details
EscStop generation

Tips & Tricks

Choose a smaller context or model before forcing a configuration that does not fit memory.
Use strict grounding when an answer must stay within the indexed sources.
Open citations and compare the extracted excerpt before relying on a generated claim.
Keep backup passphrases separate from the encrypted backup file.

Troubleshooting

ProblemWhat to check
A model cannot loadCheck compatibility, memory estimate, storage, file verification, selected backend, and another active model session.
Generation is slowUse device diagnostics, a smaller model or context, an appropriate backend, and measured rather than assumed performance.
Knowledge has no answerReview document extraction, index status, retrieval test, filters, and whether the source actually contains evidence.
A backup cannot restoreVerify the passphrase, file integrity, schema compatibility, free space, and the restore review before replacing current data.

Privacy

  • Prompts, conversations, documents, embeddings, retrieval, and generated text stay on-device.
  • Network use is limited to requested catalog/search, model download, update check, or an API endpoint you explicitly enable.
  • Mobile share inputs are temporary and external originals are never deleted.
  • App-owned workspace data and managed copies can be removed through the reviewed privacy workflow.
Your responsibility: Do not expose the optional desktop API beyond a trusted boundary or use models and documents contrary to their licenses.