joel taylor pedrós
blog

claude code for studying: an mcp server with all the exams

diagram. claude code talks to an mcp server inside a docker container on localhost, and lancedb is a folder inside the same process. the server reads and writes the markdown vault, which obsidian also reads.

my private tutor is claude code. it has read every exam from my degree, and it remembers more than i do.

studying at uni is mostly an information management problem. notes, past exams, practicals, marks, hints lecturers drop in class, and all of it scattered. two weeks before an exam, the question isn't what to study, it's what to study first.

the code is public as a template, uni-assistant-template. it's 1,846 lines of python, counting the scripts. the vault with my notes and exams is private.

the mcp server and its 19 tools

a model reads an exam well, but it doesn't know what day it is, doesn't add up weightings reliably and doesn't remember tuesday's session. for each of those gaps there's a tool in an mcp server, written in python with fastmcp. there are 19:

grouptools
timeget_current_time, get_session_duration
dashboardget_dashboard
subjects and markslist_subjects, get_subject, compute_marks_status
campaignslist_exams, get_campaign, update_campaign
loglog_progress, get_session_log
ingestprocess_ingest_folder, ingest_file, rebuild_index
searchsearch_knowledge
pdfsrender_page, export_markdown_to_pdf
gitgit_sync, git_pull

the server ships with instructions that claude code receives when it connects. the first says to call get_current_time() and then get_dashboard() at the start of every session, and the tool's documentation insists the time must never be inferred from the conversation. the docker container runs on utc by default, and on 10 june i had to add the time zone so the time would match the one at home.

compute_marks_status does the arithmetic. it reads the components of each subject's mark, with their weight, their minimum and whether they can be resat, and works out what's needed in what's left:

needed_total = passing_threshold * total_weight
already_have = current_contribution
needed_on_remaining = (needed_total - already_have) / unscored_weight

if not even a 10 (full marks) on everything left gets you to a pass, the tool says it's mathematically impossible and recommends spending the time on something else. if a component that can't be resat has fallen below its minimum, it warns that it's an overall fail.

from pdf exams to the vault

to feed it, i drop the pdfs into vault/ingest/. process_ingest_folder() lists them with a first classification that comes only from the file name. it looks for words like parcial, examen, recuperació, apunts or diapositives (midterm, exam, resit, notes, slides), in catalan, spanish and english, and for the year with a regular expression. if the name is clear, the tool says it can be ingested. if not, the agent asks me what it is before calling ingest_file().

from an exam, pymupdf extracts the text page by page. if a page has embedded images or fewer than 80 characters of text, it also renders it as a jpeg at 150 dpi. it's a crude rule, but it covers the two cases that matter, a diagram that doesn't show up in the text and a scanned exam with no text at all. claude reads images, so those pages reach it as images.

the result is an exam.md. the frontmatter gives the subject, the exam type, the year, how many pages it has and which ones are visual, and below it comes the text of each page with its images alongside. the original pdf is copied to a raw/ folder that git ignores.

on the first day, 83 exam pdfs from four subjects went in, 13 of them scanned.

a pdf's path in seven steps: the ingest folder, the classifier, pymupdf, exam.md, the 500-word chunks, the embeddings and the lancedb table.
the numbers are the ones in the code. the step in blue is the one i'd change first.

semantic search with lancedb

the vector database is lancedb, which has no server. lancedb.connect() takes a folder, and the database lives inside the mcp server's process. there's a single table, vault, with each chunk's text, the vector and nine more fields: the id, the source file, the type, the subject, the semester, the topic, the source, the quality and the date.

chunks are split by words:

def _chunk_text(text: str, chunk_size: int = 500, overlap: int = 50) -> list[str]:
    words = text.split()
    chunks = []
    i = 0
    while i < len(words):
        chunk = " ".join(words[i:i + chunk_size])
        chunks.append(chunk)
        i += chunk_size - overlap
    return chunks or [text]

each chunk goes through all-MiniLM-L6-v2, a sentence-transformers model that runs on the same machine and returns 384-dimension vectors, with no paid api involved.

search_knowledge turns the question into a vector, finds the five closest chunks and filters them by subject and by content type, which can be notes, exams, regulations or campaigns. for each result it returns 400 characters and the file path, so the agent can open the whole document if it needs to.

from a vps to my computer in a day

on 8 june, the first day, the server lived on a vps with coolify, with its own subdomain and an api key. before putting anything on it, it needed two fixes that are typical of a remote mcp server. BaseHTTPMiddleware, the starlette middleware i used for the key, buffers the whole response before sending it and broke streaming, so i swapped it for a pure asgi middleware. and fastmcp's check against dns rebinding rejected requests until i normalised the Host header.

at five in the afternoon i ingested the 83 exams. almost three hours later i moved everything to a docker container on my computer, open only on localhost:8000. the same commit computes the embeddings in batches of 32 so it doesn't run out of memory, and the project plan now says everything runs on a single machine, with no network latency and no ram limits. access from my phone through telegram was ruled out that same day.

lancedb and the local embeddings were there from day one. in the plan, lancedb is there because it doesn't need a separate process, and the model because it's local and free. after the move, the whole system is one container next to the vault.

study campaigns

a campaign is one campaign.md per exam, with a queue of past exams, newest to oldest, and for each one the exercises done and the ones left. update_campaign moves an exercise from one list to the other, and when none are left, it marks the exam as done.

the priority rule is in the readme. the default target is a 5, the pass mark, in every subject. the project plan justifies it by saying time is zero-sum, and an hour on a subject you're already passing is an hour you don't put into the one you might fail. you study with real exams from day one, and do a quick pass over everything before going deep on anything.

get_dashboard sorts the active campaigns by days left until the exam, adds the deadlines for the next 14 days and suggests starting with the most urgent. log_progress, which the agent has to call after every answer, writes an entry with the time, what was done, how it went and the minutes since the previous one. when i come back the next day, get_session_log says where i left off.

the exam patterns, which topics come up every year and which have never come up, are something claude finds by reading the exam.md files in the queue with these tools. the code only adds a heuristic to the dashboard, which counts the headings repeated in three or more exams.

markdown, git and obsidian

everything the system knows is markdown with yaml frontmatter. each subject has an INDEX.md with the components of its mark, each campaign its campaign.md and its log.md, each exam its exam.md. the server's tools, claude code and obsidian all read the same files. obsidian adds the concept graph of the whole degree, and git the history, with git_sync to commit and push from the conversation.

on 9 june i added a rule to the agent's instructions. vault first, memory second. when i give it a new fact, a date or a decision, it has to update the file in the vault before its memory, because the memory is a summary that comes from the vault, and never the other way round.

on 13 june i published the template, without the vault.

what i'd change

three things don't survive a careful read of the code.

the first is the chunk size. according to the model card, anything over 256 word pieces gets cut off, and every word is at least one piece. of each 500-word chunk, the vector sees at most the first 256. since a new chunk starts every 450 words, in a long document at least 194 of every 450 words don't make it into any vector. on top of that, the card says language: en, and the classifier looks for catalan words like apunts or recuperació.

the second is that search_knowledge reloads the model on every query. the indexing function already takes the model as a parameter, and search doesn't make use of it.

the third is that the dashboard heuristic counts headings, and every exam.md has one heading per page: ## Page 1, ## Page 2. with three exams from one subject, "page 1" already shows up as a recurring topic.

the first is the one that matters most. in a long set of notes, more than 40% of the text never makes it into a vector, and no semantic search can find it.