Specialist capability without a permanent cloud dependency.
Pestle-27B-Ternary is a medical-focused language model derived from the Qwen3.6-27B architecture and compressed for local serving. It is a public proof point for a broader Doses AI capability: making specialist models smaller, private, and practical to operate inside the institution that owns them.
8.2× smaller deployed text weights
Pestle carries 6.75 GB of deployed text weights, compared with 55.56 GB for the source Qwen3.6-27B FP16 model. The complete shipping package is one 8.48 GB runnable GGUF.
Weights, runtime, protocol, evidence.
The release pairs the runnable model with the Mortar inference runtime, deterministic generation settings, full-suite benchmark summaries, and retained evaluation evidence. Developers can inspect the operating point rather than relying on a single leaderboard number.
Measured across clinical knowledge, literature and retrieval.
Pestle is evaluated across multiple medical task families, from USMLE-style reasoning to open-ended biomedical evidence extraction and pharmaceutical retrieval.
| Benchmark | Metric | Score |
|---|---|---|
| MedQA | Accuracy | 89.79 |
| MMLU medical | Aggregate | 86.89 |
| BioASQ | Exact match | 55.94 |
| PharmaRAG | nDCG@10 | 84.62 |
| PubMedQA | Macro F1 | 62.78 |
| ChemBench | Score | 61.72 |
| MedXpertQA | Accuracy | 32.49 |
| HumanEval | pass@1 | 89.02 |
Scores and protocols follow the retained release record. See the model card for task definitions, row counts, generation settings, comparison provenance, and full evidence paths.
One local model. Several narrow, reviewable workflows.
Pestle is designed as a building block for assistive systems whose sources, prompts and outputs remain under the operator's control.
Clinical reasoning support
Draft explanations over structured medical questions and institution-approved knowledge, with professional review.
Biomedical evidence
Extract grounded spans and concise answers from supplied abstracts, policies and literature.
Private retrieval
Generate answers beside local formularies, guidelines or pharmaceutical evidence stores.
Health-tech engineering
Support coding and structured-output workflows without routing proprietary context to a public model API.
Keep the model next to the data.
Local deployment removes external inference as a structural dependency, letting organisations design access, retrieval, logging and review inside their own governance boundary.
Compress models already adapted to the institution.
Doses AI can work with specialist checkpoints and local acceptance suites to produce smaller deployment artifacts without making a third-party inference endpoint mandatory.
Keep therapeutic-area models and evidence private.
Compress medical-information, literature, safety, regulatory or retrieval models for operation near proprietary corpora and governed internal workflows.