Skip to content
View jrcastro2's full-sized avatar
  • Geneva, Switzerland

Block or report jrcastro2

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jrcastro2/README.md

Javier Romero Castro

Backend software engineer based in Geneva. I've spent the last 5+ years at CERN building and maintaining scientific information platforms — primarily CDS and its migration to InvenioRDM, an open-source research data management platform used by institutions worldwide.

I work mostly in Python, with a focus on REST APIs, search infrastructure, and data workflows.


Stack

Backend — Python, Flask, PostgreSQL, OpenSearch/Elasticsearch, Celery, REST APIs
Frontend — React, TypeScript
Infrastructure — Docker, Kubernetes/OpenShift, CI/CD
Domain — research data, open science, metadata standards, scientific repositories


Projects

A metadata assistant for scientific papers built to learn LLM tool calling and agentic workflows hands-on. Uses the Anthropic API and queries the INSPIRE HEP database. The model decides which tools to call and in what order — search, citation lookup, metadata fetch — without that sequence being scripted.

An MCP (Model Context Protocol) server that exposes InvenioRDM record management as tools an LLM can call. Lets an AI assistant create drafts, set metadata, upload files, and publish records on a research repository through natural language.

A retrieval-augmented generation (RAG) system for question-answering over scientific paper abstracts. Built in stages to understand each component: hybrid search (dense vector + BM25 fused with RRF), cross-encoder reranking with a relevance threshold, contextual RAG enrichment, and grounded generation with Claude. Indexes real papers from the arXiv API.

Two LLM modules for processing scientific paper abstracts: a metadata extractor that returns reliable structured JSON (title, authors, identifiers) from clean or messy input, and a classifier that assigns arXiv-style categories using both zero-shot and few-shot prompting — with a side-by-side comparison to show where each approach wins.


Open source

Most of my open-source work is through contributions to the InvenioRDM ecosystem — the framework behind Zenodo, CDS, and a number of institutional repositories.


Contact

jrcastro9515@gmail.com

Pinned Loading

  1. invenio-mcp invenio-mcp Public

    MCP server for InvenioRDM — lets an LLM create and publish records via tool calling

    Python 2

  2. paper-assistant paper-assistant Public

    Agentic metadata assistant for scientific papers — LLM tool calling with the Anthropic API and INSPIRE HEP

    Python

  3. rag-paper-assistant rag-paper-assistant Public

    RAG pipeline for question-answering over scientific paper abstracts. Hybrid search (vector + BM25), cross-encoder reranking, contextual RAG, and Claude for generation.

    Python

  4. paper-metadata paper-metadata Public

    LLM modules for scientific paper processing: structured metadata extraction and zero-shot/few-shot classification.

    Python