Veridion
An image-embedding-based copyright-risk detection and review-support system that separates text, narrative and character-identity signals, then assembles evidence with ontology and provenance
- requirements, architecture, backend, frontend and ontology implementation
- 100% SOLO
- separate signals for file reuse and redrawn-character identity
- SSCD · CCIP
- scale-aware cut-points with abstention beyond measured evidence
- 1:N · FPIR
- traceable, retractable evidence in named graphs
- RDF · PROV-O
What it does
- Defined the product requirements and evidence boundaries, then single-handedly delivered the system architecture, backend, frontend and ontology implementation.
- SSCD near-duplicate search, CCIP character identity, Korean verbatim and semantic retrieval, and narrative, speech-style and persona rubrics remain separate evidence boundaries.
- Fixed cosine verdicts are replaced by cut-points derived from an FPIR target and the actual comparison count; beyond empirical resolution, the system reports similarity but withholds the label.
- Whole-frame and cropped-query impostor cohorts are calibrated separately; per-region retrieval and round-robin merging stop identities with more references from monopolising top-k in multi-character frames.
- The LLM cannot write Turtle directly: it emits a fixed JSON contract, the server builds allowlisted triples, and model-derived claims are isolated by extractor in named graphs that can be withdrawn as a unit.
- After measuring identity collapse under stylisation and a regression from visual-trait fusion, traits were removed from automatic verdict scoring and retained as explanation for human review.
Measurements behind these claims
- Veridion open-set calibration: scale-aware thresholds and measured abstention Recasts character retrieval as 1:N open-set identification rather than a fixed cosine cut-off. It derives a false-positive budget from gallery and region count, abstains beyond empirical resolution, calibrates whole-frame and crop cohorts separately, and merges unequal galleries fairly. 1:N · FPIR the cut-point is a function of gallery size, region count and review budget—not a model constant
Stack
Python FastAPI PyTorch ONNX Runtime FAISS Next.js TypeScript LLM Agents Oxigraph RDF/SPARQL PROV-O PostgreSQL
Skills demonstrated
-
AI Agent Systems
Agent loops · tool contract design · MCP client and server (OAuth 2.1 · PKCE) · GraphRAG · Personalized PageRank · CSLS · Offline evaluation harnesses · Computer use · approval gates · threat modelling
-
Ontology & Knowledge Graphs
RDF / W3C quads · named graphs · SPARQL 1.1 · Oxigraph · RocksDB · SKOS · Dublin Core · schema.org · PROV-O · GeoSPARQL · Entity resolution (union-find · embeddings · LLM adjudication) · Louvain / Leiden community detection
-
Product & Full-stack
TypeScript · JavaScript · Python · C/C++ · C# · Java · Next.js · Nuxt.js · Vue.js · React · React Native · Expo · PostgreSQL · Supabase · AWS · GCP · Docker · UX design for ERP, CRM and LMS operational screens
-
Research Methodology & Benchmarking
Pre-registered hypotheses, rejection thresholds and the commit hash at registration · Paired statistical tests implemented directly (McNemar · Fisher · Wilcoxon · Holm-Bonferroni) · Same-arm controls to establish the noise floor; pooled replicates to test whether significance reproduces · Blinded LLM judge panels, false-negative controls, and deterministic scoring with no judge · Adversarial refutation rounds; null results and failures to reproduce on the public record; do-not-re-propose lists · Public benchmarks measured first-hand: SWE-bench · MuSiQue · DP-Bench · OmniDocBench