Text classification & decisions
JevAlt
Route a support ticket, apply a policy, or choose an action. Three 4B language models return decisions with probabilities through the Jev API.
96.8% held-out accuracy · Kev-4B 87.1%
94.7% held-out accuracy · Kev-4B 84.7%
92.0% held-out accuracy · Kev-4B 81.1%
Held-out tests share the training data pipeline. These scores measure the target decision task. Q4_K_M uses about 3 GB RAM at 4k context with the JevAlt server. Refit calibration on your own data and keep authorization outside the model.
Local AI agents & tool use
Tholos-2B
A 2B model that chooses tools for tables, notes and tasks on your own machine.
137/160 tasks
Tholos-Bench, Q4_K_M, llama.cpp JSON schema, Kaggle T4. MiniCPM5-2B: 112/160. Only Tholos-2B was fine-tuned for this task format.
Q4_K_M file: 1.56 GB. Use Ollama or llama.cpp. Validate tool calls and retain approvals for consequential actions.
Explainable deepfake detection
xdfdet
Eight EfficientNet-B4 detectors with Grad-CAM maps showing which facial regions influence a prediction.
0.8981AUC on FaceForensics++
Best released checkpoint, aug-cutout-black. Baseline: 0.8684. Each checkpoint used its own random split. On 398 unseen DFDC videos, AUC drops to 0.60 to 0.66.
Weights are CC BY-NC 4.0 under FaceForensics++ terms; code is MIT. A score alone cannot establish a video’s authenticity.