Case study
AniReco
Personalized anime recommendations from your AniList
- Hybrid ranker
- 5 signals
- Retrieval
- 3 vectors
- Indexed catalog
- 50k+ anime
- Hot-path fix
- 4 min → sub-10s
- Product surface
- OAuth + billing
- Outcome
- Archived (economics)
- Python
- FastAPI
- React
- ChromaDB
- MongoDB
- Redis
- Gemini
90-second story
Popularity lists are useless for taste. AniList already holds yours. I built a hybrid recommender (embeddings + IDF tags + franchise graph policy) and shipped the full product: OAuth, subscriptions, streaming API. Naive vector search drowned in common tags; score-relative profiling and dislike vectors fixed quality. Production taught me the harder lesson: cold recommendation paths hit Gemini, Chroma, and AniList on every miss. Usage scaled with curiosity; revenue didn't. I archived it instead of bleeding.
The story
AniList knows what you watched, rated, and dropped. Popular “top anime” lists don’t know you.
The problem: Generic rankings ignore personal taste. Pure embedding nearest-neighbors ignore that tags like “Comedy” and “Male Protagonist” appear on 70%+ of the catalog. Without rarity weighting, those tags dominate every profile.
Constraints I designed around:
- AniList rate limits and OAuth scope. The server should not scrape a user’s full list on every request.
- Chroma on SQLite: concurrent reads OK, writes must serialize.
- Gemini embedding cost scales with catalog growth and cache misses.
- Anime isn’t flat metadata. Sequels, prequels, and franchise order matter for “what to watch next.”
Approach: what I considered
- Popularity + genre filtersRejected
Fast to build, zero personalization. Fails the core promise.
- Embeddings-only k-NNRejected
Common tags swamped profiles; no dislike signal; franchise chaos (recommending part 3 of 3).
- Hybrid: vectors + IDF tags + graph gate + MMRChosen
Each signal fixes a specific failure mode. Tunable weights; explainable reasons.
I shipped the full stack (React SPA, FastAPI, auth, PayPal subscriptions, GA4 Measurement Protocol), then archived when marginal cost per fresh compute outran monetisation. The code stays in the repo; the domain now serves this portfolio.
How it works
In one sentence: Your browser sends your AniList list to the server; the server figures out your taste, searches a big catalog of pre-indexed anime, scores and filters candidates, then streams back IDs for the UI to display.
The flow in plain language
This is what actually happens when you ask for recommendations:
- You enter an AniList username. The site loads your list in the browser: what you completed, dropped, are watching, and how you scored each title.
- Your browser sends that list to the server. We do not pull your full list from AniList again on every refresh. That keeps us inside AniList rate limits and avoids extra OAuth scope on the backend.
- The server builds a taste profile from your list. It learns what you tend to like (high ratings vs your average, not vs everyone), what you tend to dislike (drops and low scores), and which tags/genres show up on the shows you loved vs the ones you didn’t.
- It searches a catalog of tens of thousands of anime. Each title was already embedded with Gemini and stored in Chroma with tags, genres, year, and AniList score. The search is “find anime whose embedding and metadata look like this person’s taste.”
- Three searches, then merge. One search targets shows you clearly loved, one your recent watches, one your overall profile. We combine the results and drop anything already on your list.
- Every candidate gets a number. The score mixes: similarity to your taste vector, match on tags you like (weighted so rare tags matter more), penalty for tags you dislike, the show’s AniList rating, and how new it is.
- We diversify the list. Without this step, the top 10 would often be the same genre and tag cluster. MMR spreads picks so the list feels varied.
- Franchise rules run on each slot. If a candidate is part of a series you’ve started, we may replace it with the right entry point (first unwatched season) or the next sequel after you finished everything, or drop the slot if it would be a bad recommendation.
- Results stream to the browser. The server sends anime IDs and scores first (small payload). The browser fetches titles, covers, and links from AniList for display.
If that makes sense, the diagram below shows where each step runs. The deep dives later explain how profiling, IDF, scoring, and the franchise gate work inside steps 3-8.
Client
React SPA
Fetches AniList list
SSE consumer
Hydrates media client-side
API
POST /compute-stream
Profiler → Engine → Gate
Data & models
ChromaDB
Embeddings + metadata
MongoDB
Lists, relations, identity
Redis
Cache, quotas, embed queue
Off hot path
Gemini
Indexing & queue worker only
Why this diagram: The browser fetches the user's list; the server profiles and scores. That avoids OAuth scope fights and AniList rate limits on every refresh.
1. List arrives from browser
Your AniList entries + settings (sequel preference, filters, etc.).
2. Taste profile
Like/dislike vectors and tag/genre preferences from your ratings and list status.
3. Find candidates in Chroma
Three vector searches, merged. Already-on-your-list titles excluded.
4. Score each candidate
Similarity + tags + genres + AniList score + recency, minus avoid-similar-to-dislikes.
5. Diversify (MMR)
Reduce duplicate genres/tags in the final list.
6. Franchise gate
Fix series order: start at the beginning, suggest sequels when appropriate, or skip.
7. Stream IDs to browser
Client loads posters and titles from AniList.
Taste profile
Before retrieval, UserProfiler turns list rows into vectors and structured preferences.
Status weights encode how much each list state matters. Completed and repeating titles count most; dropped counts negative:
STATUS_WEIGHTS = {
"COMPLETED": 1.0,
"REPEATING": 1.2,
"WATCHING": 0.6,
"PAUSED": 0.2,
"PLANNING": 0.05,
"DROPPED": -0.8,
}Score-relative weighting: A 5/10 from someone whose mean is 7.5 is displeasure, not neutral approval. The old 0.5 + score/10 formula made every scored entry positive.
Entry weight = status × score deviation multiplier
Dislike vector: Not just DROPPED. Anything scored ≤ mean − 1.5 enters the dislike pool. That catches “finished but didn’t enjoy” without requiring harsh 1-4 scores.
Why IDF
Tags like “Male Protagonist” or “Comedy” appear on most anime. Without rarity weighting they dominate every user’s profile and drown out distinctive taste (“Psychological”, “Band”, “Seinen”).
IDFService: normalized inverse document frequency
Tag weighting options
- Raw tag countsRejected
Common tags always win; profiles look identical.
- Manual blocklist of generic tagsConsidered
Works until AniList adds tags; brittle to maintain.
- IDF from full Chroma catalogChosen
Data-driven rarity; auto-updates when catalog grows.
After aggregation, a score adjustment per tag uses avg_score for that tag vs the user’s mean. Tags they consistently underrate get dampened even if frequent.
Hybrid score
Each candidate gets a weighted sum of five signals. Tag and genre contributions live in [-1, 1]. Disliked features drag the total down.
RecommendationEngine.BASE_WEIGHTS
Weights adapt slightly when the profile is sparse (cold start) vs rich: more vector weight early, more tag/genre weight when preferences are well-estimated.
Multi-vector retrieval
One centroid embedding of “everything you liked” blurs strong taste with recent experiments and long-tail completions.
Three query vectors:
- Strong preference: top-rated entries only (≥ mean + offsets). “What you unambiguously love.”
- Recent preference: last ~18 months of signal. “What you’re into now.”
- General preference: full weighted profile. “Broad taste baseline.”
Chroma returns candidates per query; the engine merges, dedupes, and excludes anything already on the user’s list upstream.
Franchise gate
Vector search doesn’t know that Fate/stay night: Unlimited Blade Works is mid-franchise. The franchise gate is a graph policy, not a checkbox filter.
Component: BFS from candidate C over AniList relations (SEQUEL, PREQUEL, PARENT, SPIN_OFF, …) cached in Mongo.
Rules (simplified):
| Situation | Sequel OFF | Sequel ON |
|---|---|---|
| Any franchise node on user’s list | Drop slot | Drop if any on-list status ≠ COMPLETED |
| Nothing on list | Recommend earliest show in component | Same |
| All on-list nodes COMPLETED | (none) | Recommend best-scored direct SEQUEL not on list |
The gate runs after MMR but over-fetches candidates beforehand (FRANCHISE_GATE_MMR_FACTOR = 3) because many slots get dropped or swapped.
MMR diversity
Top-k by score alone returns ten psychologically similar thrillers. Maximal Marginal Relevance re-ranks with a penalty for genre and tag overlap with picks already chosen.
MMR selection
λ comes from user options (diversity_lambda). Lower λ means more diverse, less strictly “best match.”
Hot path vs money path
Production hit a wall: recommendation requests stalled 4+ minutes. Root cause stack:
- Single-thread Chroma pool blocking reads behind writes
- Gemini retries with long sleeps inside the request path
- Per-slot franchise lookups fanning out on the same pool
Latency fix
- Widen the single Chroma thread poolRejected
SQLite-backed Chroma still serializes writes; risk write conflicts.
- Dual-pool Chroma (multi read + single write)Chosen
Parallel reads; safe serialized upserts.
- Redis embedding queue + background workerChosen
Gemini indexing off the hot path; queue can lag without blocking users.
After the fix: hot path sub-10s with graceful time budgets; cold paths still expensive.
The economics lesson: Cache hits were fine. Every cache miss on a fresh list could touch embedding API, Chroma upserts, AniList relations, and Redis. Cost scaled with curiosity, not revenue. See the unit economics note for the full breakdown.
Product around the model
This was a product, not a notebook:
- Identity-first users:
users/identities/entitlementsso Google, GitHub, and email auth share one billing anchor. - SSE slim payload: IDs + scores stream first; client hydrates media (smaller responses, faster TTFB).
- Redis: recommendation cache keyed by list fingerprint + options; rate limits and quota tiers.
- GA4 Measurement Protocol: single server-side pipeline; Ads conversions imported from GA4 key events.
What I took away
Technically: Hybrid beats any single signal for anime taste. IDF and score-relative profiling are non-optional when metadata is tag-heavy. Graph policies belong in the product layer, not as an afterthought filter.
Operationally: Profile and cache before you optimize the model. Measure cache miss cost, not average latency.
Economically: For consumer AI, margin per request is the moat, or the kill switch. I archived AniReco instead of pretending subscriptions would catch up.
What I’d do differently:
- Model unit economics before indexing 50k titles.
- Separate “demo” (cached replays) from “daily driver” (unlimited fresh computes).
- Ship the writeup earlier. The system deserved this explanation while it was live.
For the economics deep-dive, read When the unit economics don’t work.