Software Engineer

🏒 Location: Pisa, Italy (CET) Remote: Yes, remote only Willing to relocate: No Technologies: Python, LLM/agent orchestration, local embeddings, retrieval evaluation, AWS (Lambda/SQS/EventBridge), Terraform, Django, PostgreSQL, LightGBM/PyTorch, NLP/Transformers RΓ©sumΓ©/CV: linkedin.com/in/vslovik Code: github.com/vslovik/fenix β€” local-embedding search whose relevance is actually measured: labelled control probes, blinded human ranking, precision@k. No API keys. Email: valeriya.slovikovskaya@gmail.com Software architect, 15+ years in production systems, almost entirely startups and internal startups β€” fintech, e-commerce, pharma, publishing. The work I get pulled into is the recurring startup problem: a service shipped fast under launch pressure, without adequate tests, that later has to be made reliable without being stopped. Incident response, re-architecture, and the release discipline that keeps it from happening again. Most recently that has meant a regulated UK consumer-credit platform β€” loan servicing, arrears, forbearance, statutory breathing space, and early-settlement calculations written against consumer-credit legislation. Regulation as code, behind a test suite larger than the production codebase. I've done that in all three configurations: taking a core system from problem statement to release, leading the team that carried it (1 to 7 engineers in ten months), and now doing the same work again with agentic tooling covering what the team used to. On the data side: a LightGBM acquisition model over a 38M-row base β€” 0.77 test AUC, 8x lift in the top 1% β€” scoring 2.9M households for a live campaign. The part I'd rather be judged on is what happened next: I found a validation-set defect in my own pipeline (early stopping on the test split), quantified its effect across every figure I had already reported, restated them, and added a pure-noise regression test that pins the model to chance when fed random features β€” so that class of leak cannot come back quietly. NLP is hands-on rather than API-deep: my degree thesis fine-tuned BERT, RoBERTa and XLNet to state of the art on the FNC-1 stance-detection benchmark, published at LREC 2020. Building on my own time: github.com/vslovik/fenix β€” it ranks an incoming stream against a free-text description of what you're looking for, and answers questions over the same corpus with citations back to source chunks. Ollama embeddings, sqlite-vec, no API keys. The part worth looking at is the evaluation: the ranking anchor is scored against a labelled probe set with a deliberate control group of things I don't want, and live results are rated blind β€” scores hidden, order shuffled β€” so the human judgement stays independent of the ranking it is judging. Doing that produced a measured finding I did not expect: an embedding has no notion of negation, so naming a technology in order to reject it moves the anchor toward it. Numbers and method in lessons/embedding-anchors.md. Also a tool-calling agent that turns unstructured regulatory text into a deterministic calculation pipeline β€” the model does the extraction, a deterministic engine does the arithmetic. Looking for agentic AI/LLM engineering, LLM evaluation and observability, AI integration, or software architecture. Founding-engineer shape suits me β€” early employee, not co-founder, but early enough to be in the room where the work gets defined. Direct with the company that owns the product: not consultancy, not agency placement, not a body on someone else's engagement. Β· company page
πŸ“ Remote
via HNWhoIsHiring
🏷 Hn Whos Hiring
View discussion on Hacker News β†—

Location: Pisa, Italy (CET)
Remote: Yes, remote only
Willing to relocate: No
Technologies: Python, LLM/agent orchestration, local embeddings, retrieval evaluation, AWS (Lambda/SQS/EventBridge), Terraform, Django, PostgreSQL, LightGBM/PyTorch, NLP/Transformers
RΓ©sumΓ©/CV: linkedin.com/in/vslovik
Code: github.com/vslovik/fenix β€” local-embedding search whose relevance is actually measured: labelled control probes, blinded human ranking, precision@k. No API keys.
Email: valeriya.slovikovskaya@gmail.com
Software architect, 15+ years in production systems, almost entirely startups and internal startups β€” fintech, e-commerce, pharma, publishing.
The work I get pulled into is the recurring startup problem: a service shipped fast under launch pressure, without adequate tests, that later has to be made reliable without being stopped.

Incident response, re-architecture, and the release discipline that keeps it from happening again. Most recently that has meant a regulated UK consumer-credit platform β€” loan servicing, arrears, forbearance, statutory breathing space, and early-settlement calculations written against consumer-credit legislation. Regulation as code, behind a test suite larger than the production codebase.
I've done that in all three configurations: taking a core system from problem statement to release, leading the team that carried it (1 to 7 engineers in ten months), and now doing the same work again with agentic tooling covering what the team used to.
On the data side: a LightGBM acquisition model over a 38M-row base β€” 0.77 test AUC, 8x lift in the top 1% β€” scoring 2.9M households for a live campaign.

The part I'd rather be judged on is what happened next: I found a validation-set defect in my own pipeline (early stopping on the test split), quantified its effect across every figure I had already reported, restated them, and added a pure-noise regression test that pins the model to chance when fed random features β€” so that class of leak cannot come back quietly. NLP is hands-on rather than API-deep: my degree thesis fine-tuned BERT, RoBERTa and XLNet to state of the art on the FNC-1 stance-detection benchmark, published at LREC 2020.
Building on my own time: github.com/vslovik/fenix β€” it ranks an incoming stream against a free-text description of what you're looking for, and answers questions over the same corpus with citations back to source chunks. Ollama embeddings, sqlite-vec, no API keys.

The part worth looking at is the evaluation: the ranking anchor is scored against a labelled probe set with a deliberate control group of things I don't want, and live results are rated blind β€” scores hidden, order shuffled β€” so the human judgement stays independent of the ranking it is judging. Doing that produced a measured finding I did not expect: an embedding has no notion of negation, so naming a technology in order to reject it moves the anchor toward it. Numbers and method in lessons/embedding-anchors.md.
Also a tool-calling agent that turns unstructured regulatory text into a deterministic calculation pipeline β€” the model does the extraction, a deterministic engine does the arithmetic.
Looking for agentic AI/LLM engineering, LLM evaluation and observability, AI integration, or software architecture.

Founding-engineer shape suits me β€” early employee, not co-founder, but early enough to be in the room where the work gets defined. Direct with the company that owns the product: not consultancy, not agency placement, not a body on someone else's engagement.

← All remote jobs

Comparing Software Engineer pay and openings β€” the live median is $129k?All remote Software Engineer jobs β†’Software Engineer salary data β†’
Get new remote jobs like this by email
Daily email, only when there's something new. One click to stop.

Get remote jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you