ClipMaker — on-device AI model credits
==========================================
All models below run 100% on your PC. No cloud, no accounts, nothing uploaded.

BUNDLED MODELS (ship inside the ClipMaker zip, under models/)
--------------------------------------------------------------

1. transnetv2.onnx — shot boundary detection (split planning)
   Source: TransNet V2 ("TransNet V2: An effective deep network architecture
   for fast shot transition detection", Song et al.). Community ONNX
   conversion of the official implementation.
   License: MIT (as stated on the Hugging Face model card).
   SHA-256: c4d54a682bace32f25136ef83ca2c9d403e8f8193775efeb995172a0d95a8e0c

2. rf-detr-nano.onnx — subject detection & labeling (replaces YOLOv8n in v1.3.0)
   Source: Roboflow RF-DETR nano, ONNX export (Apache-2.0).
   License: Apache-2.0.
   SHA-256: 3fcbba0f68bad4939fdf1c38f432783b95691e2869af3be369780aa5be67abb2

3. lutwithbgrid.onnx — learned color enhancement (AI Grade)
   Source: "Learning Image-adaptive 3D Lookup Tables for High Performance
   Photo Enhancement" (Zeng et al.); FiveK-trained weights, ONNX conversion.
   License: Apache-2.0 (LICENSE file verified in the source repo).
   SHA-256: a07428e5cfa9c3db8e8ebdfd8a90ad7240e4d31a2260b45468ba17ebd3ebe561

4. face_detection_yunet_2023mar.onnx — face detection (YuNet, 2023mar)
   Source: opencv_zoo model zoo (OpenCV team),
   https://github.com/opencv/opencv_zoo/tree/main/models/face_detection_yunet
   License: MIT (as stated in the opencv_zoo repo for this model).
   SHA-256: 8f2383e4dd3cfbb4553ea8718107fc0423210dc964f9f4280604804ed2552fa4

5. face_recognition_sface_2021dec.onnx — face recognition (SFace, 128-dim)
   Source: opencv_zoo model zoo (OpenCV team),
   https://github.com/opencv/opencv_zoo/tree/main/models/face_recognition_sface
   License: Apache-2.0 (as stated in the opencv_zoo repo for this model).
   SHA-256: 0ba9fbfa01b5270c96627c4ef784da859931e02f04419c829e83484087c34e79

ON-DEMAND MODELS (downloaded by ClipMaker on first use with your approval;
workspace copies are kept here for tests — see SHA256SUMS.txt)
----------------------------------------------------------------

6. RealESRGAN_x4plus.fp16.onnx — AI 4x super-resolution (upscale)
   Source: Real-ESRGAN weights by xinntao, fp16 ONNX conversion by tamnvcc:
   https://huggingface.co/tamnvcc/RealESRGAN-onnx/resolve/main/onnx/RealESRGAN_x4plus.fp16.onnx
   License: BSD-3-Clause (Real-ESRGAN's license).
   SHA-256: 0a06c68f463a14bf5563b78d77d61ba4394024e148383c4308d6d3783eac2dc5

7. Whisper base.en (int8 quantized) — speech transcription
   encoder_model_quantized.onnx / decoder_model_quantized.onnx /
   decoder_with_past_model_quantized.onnx + vocab.json + merges.txt
   Source: OpenAI whisper-base.en weights, ONNX int8 conversions by Xenova:
   https://huggingface.co/Xenova/whisper-base.en/tree/main/onnx
   Tokenizer (vocab.json, merges.txt) from OpenAI's repo:
   https://huggingface.co/openai/whisper-base.en/resolve/main/vocab.json
   https://huggingface.co/openai/whisper-base.en/resolve/main/merges.txt
   License: MIT (as stated on the model cards/repos).
   SHA-256:
     encoder_model_quantized.onnx:
       d0d4e59e2842617b39787cece73d7e8f76f99b1697d3386c0e682eca2269f4a1
     decoder_model_quantized.onnx:
       295b8ae738ab8b39612e7ee1ee8cef31e865733f3ecf33d5d319a11df66bf8d0
     decoder_with_past_model_quantized.onnx:
       9d57a5d2b6c2e247c2e26818a937fdd9338ac90a961bcc6908eff422b45cf6ef
     vocab.json:
       3ba3c3109ff33976c4bd966589c11ee14fcaa1f4c9e5e154c2ed7f99d80709e7
     merges.txt:
       1ce1664773c50f3e0cc8842619a93edc4624525b728b188a9e0be33b7726adc5

8. fastdvdnet_s25.onnx — temporal video denoise (AI denoise, sigma=25)
   Source: m-tassano/FastDVDnet weights, community ONNX conversion from the
   npuscale project (v1.3 release):
   https://github.com/ikeno-web/npuscale/releases/download/v1.3/fastdvdnet_s25.onnx
   Note: this is a small third-party conversion project, not an official
   release; the conversion was not independently audited. Upstream weights
   are m-tassano/FastDVDnet. Kept for evaluation; ClipMaker's default
   denoise path is ffmpeg's hqdn3d.
   SHA-256: eb075f1c69fb36ab7db10e5f568ba5b2873692047ebfaa110c1e0ab77ba7b939

FFmpeg (ffmpeg.exe) is bundled under its own licenses (GPL/LGPL — see the
FFmpeg project's LICENSE files). ONNX Runtime is MIT licensed.

9. SeedVR2 3B (FP8) — Pro AI upscale engine (v1.2, on-demand, not bundled)
   Weights source: numz/SeedVR2_comfyUI (Hugging Face), pinned revision
   09ced71023636e9bc8cdf9cdecfb2625d1e691e8 (Apache-2.0):
   https://huggingface.co/numz/SeedVR2_comfyUI
     seedvr2_ema_3b_fp8_e4m3fn.safetensors (3,391,544,696 bytes, FP8 DiT):
       3bf1e43ebedd570e7e7a0b1b60d6a02e105978f505c8128a241cde99a8240cff
     ema_vae_fp16.safetensors (501,324,814 bytes, fp16 VAE):
       20678548f420d98d26f11442d3528f8b8c94e57ee046ef93dbb7633da8612ca1
   Note: FP8 was chosen because it fits a 16 GB VRAM GPU; the upstream
   FP16 3B pack needs ~17 GB peak and would OOM. NVFP4 packs are a future
   upgrade path (no pinned public NVFP4 3B pack existed at pin time).
   Sidecar driver: ussoewwin/ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT,
   pinned commit 3bf2985bfa15741024aa32c737e6f2d7d372de95 (Apache-2.0),
   run standalone as inference_cli.py (its ComfyUI imports are lazy).
   SHA-256 of the driver zip:
     f239a4683893baca315703532818a76dcf38e2e505cf0c71c29c88311c3f73b2
   Pins verified 2026-09-30 from three independent sources; see
   src/ClipMaker/Models/SeedVr2Catalog.cs. Upstream SeedVR2 project:
   https://github.com/numz/SeedVR2

10. rife_v4.25_v2.onnx — AI frame interpolation (RIFE v4.25, slow-motion /
    frame-rate doubling; the Topaz Chronos equivalent)
   Source: Hugging Face notaneimu/onnx-image-models (community re-host of the
   vs-mlrt ONNX conversion):
   https://huggingface.co/notaneimu/onnx-image-models/resolve/main/rife_v4.25_v2.onnx
   Upstream weights: Practical-RIFE v4.25 by Zhewei Huang (hzwer/Practical-RIFE),
   MIT licensed (RIFE: Real-Time Intermediate Flow Estimation for Video Frame
   Interpolation, ECCV 2022; (c) Megvii Inc.). ONNX export by the vs-mlrt
   project (AmusementClub/vs-mlrt, "v2" implementation with internal padding).
   Verified 2026-09-30 by direct download + empirical ONNX Runtime run:
   single input "input" (1,7,H,W) float32 = [img0 RGB, img1 RGB, timestep plane],
   single output "output" (1,3,H,W) float32, values ~[0,1]; any H/W accepted
   (internal padding); timestep honored (t=0 -> img0, t=1 -> img1, t=0.5 ->
   midpoint). Third-party conversion, not independently audited beyond the
   smoke test above; quality degrades on fast motion / occlusions / scene cuts.
   SHA-256: 65c57a5e4abb17ad67faf35054291ac53affab0506d509d2e66698f6ecd75584
   Size: 22,748,049 bytes. On-demand download (not bundled); used with the
   DirectML execution provider first, CPU fallback.

ON-DEMAND MODELS (v1.3.0, downloaded once on first use)
--------------------------------------------------------

A. rife_v4.26.onnx — AI frame interpolation (upgraded v4.25 -> v4.26)
   Source: Practical-RIFE v4.26 (hzwer), self-exported to ONNX in-build.
   License: MIT.
   SHA-256: 98a5649eb2778c5342d9d0609c75a6eec7720a94f5f1dfb24c5dc6eafcc67ca0

B. gfpgan/GFPGANv1.4.pth — blind face restoration
   Source: TencentARC GFPGAN v1.4 (release v1.3.4).
   License: Apache-2.0.
   SHA-256: e2cd4703ab14f4d01fd1383a8a8b266f9a5833dacee8e6a79d3bf21a1b6be5ad

C. gfpgan/detection_Resnet50_Final.pth — facexlib RetinaFace detector
   Source: xinntao/facexlib release v0.1.0.
   License: Apache-2.0.
   SHA-256: 6d1de9c2944f2ccddca5f5e010ea5ae64a39845a86311af6fdf30841b0a5a16d

D. gfpgan/parsing_parsenet.pth — facexlib face parsing network
   Source: xinntao/facexlib release v0.2.2.
   License: Apache-2.0.
   SHA-256: 3d558d8d0e42c20224f13cf5a29c79eba2d59913419f945545d8cf7b72920de2

E. flashvsr/* — FlashVSR v1.1 Tiny "Ultra-fast" upscaler (4 files, ~6.9 GB)
   Source: OpenImagingLab/FlashVSR (Hugging Face: JunhaoZhuang/FlashVSR).
   License: Apache-2.0.
   SHA-256: PENDING (too large to hash in the build VM; sizes verified).
