I build AI agents, and write papers about why they break.

Forward Deployed Engineer at Wonderful. Founder of Synkrasis Labs. I ship agents into enterprises and study what happens when they meet environments we don’t control.

Also, I really like the sea.

Athens, Greece
Panos Michelakis paddling a SUP board on calm water off the coast
Field research.
Research

Selected publications

NeurIPS - LAW2025

CORE: Full-Path Evaluation of LLM Agents Beyond Final State

P. Michelakis, Y. Hadjiyiannis, D. Stamoulis

A framework built on finite automata, with five metrics that score an agent's entire execution path, not just whether the final answer happens to be correct.

arxiv.org/abs/2509.20998 ↗
ICML - AIWILD2026

Full-Season Agent Evaluation in Soybean Farm Operations under Real-World Agricultural Process Dynamics

A. Qu, P. Michelakis, Y. Hadjiyianni, F. Li, J. Jiang, D. Stamoulis, J. Liu

Benchmarks nine agent methodologies on a real soybean farm and finds that agents trail expert human yields by 34% without expert context; long horizon performance hinges on it.

openreview.net ↗
SIGSPATIAL - Applications2026

Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations

A. Qu, P. Michelakis, L. Han, Y. Hadjiyianni, K. Ouyang, K. Siskos, F. Li, R. Meng, J. Jiang, D. Stamoulis, J. Liu

A full-stack agent system deployed on an operating 64-ridge soybean farm, built on an “everything is an event” execution engine. Benchmarks nine agent controllers over a hundred full-season scenarios with a yield-aligned spatiotemporal correctness metric.

arxiv.org/pdf/2609.00106 ↗
PRICAI2026

Position: We Should Evaluate Agentic Memory as Inference, not as a Tape Recorder

Y. Hadjiyianni, P. Michelakis, E. Neofotistos, J. Jiang, D. Stamoulis, J. Liu

Reframes memory evaluation as inference over episodic evidence rather than recall. Interventions separate models that look identical on accuracy: citation precision spans 0.54 to 1.00 at the same task score.

arXiv link pending
MLSys - YPS2026

QPU-first ML kernels for Raspberry Pi 5

P. Michelakis et al.

A compact ML runtime with integer matmul and neural network kernels targeting the Pi 5's VideoCore VII GPU for efficient edge inference.

arxiv.org/abs/2606.09905 ↗

All agents are wrong, but some are useful.

after George Box
Work

Experience

2026 to now
Wonderful · Athens (Hybrid)
  1. Pod Tech Lead
    Sep 2026 to now

    Technical lead for a six person pod delivering agent systems into banking and insurance. I set the architecture and agent design standards across the vertical's accounts, review the team's work, and stay hands on in the deployments that need it.

  2. Forward Deployed Engineer
    Jan 2026 to Sep 2026

    Designed and deployed production AI agent systems for enterprise clients: voice agents, back office automation, and workflow orchestration, wiring LLM agents into CRM, ERP, telephony, and external APIs on AWS. Ran engagements end to end, from scoping with executives to production rollout, including work on deployments worth millions.

2024 to now
Founder & Research Lead
Synkrasis Labs · Athens

Founded an independent research group focused on deploying AI agents in real environments and evaluating what they do beyond the final answer. Our work covers full-path evaluation, long-horizon farm operations, agent memory, and edge inference.

See our research ↗
2025 to 2026
Palantir Foundry Data Engineer
D ONE · Zurich (Remote)

Built the enterprise Data Quality Framework for a leading Swiss reinsurer: PySpark pipelines, a scalable Ontology architecture, a custom PySpark library for advanced data quality checks, and automated monitoring via TypeScript functions. Also shipped scheduling tool UI features using Vertex Graphs and Workshop.

2023 to 2025
Founding Machine Learning Engineer
Vino AI · Athens (Hybrid)

Cofounded an AI hospitality startup. Designed and productionised a real time KNN recommendation engine that served personalised drink suggestions from taste preferences collected via QR menus, plus the full backend and cloud architecture. Launched across multiple venues.

About

How I work

I like the gap between research and production. It's usually where the interesting problems are.

A lot of my research is just being suspicious of agents: measuring the whole path one takes, instead of whether it fluked the right answer at the end. That suspicion bleeds into what I build. I'd rather know how something fails than pretend it won't.

Day to day that's turning vague requirements into systems that hold up, and switching between talking to an executive and a compiler without too much whiplash. I've done the founding thing, the enterprise data thing, and the research thing, and I'm not in a hurry to pick one.

When I have something worth saying, it ends up on Medium: agents, physics, math and my occasional poetry.

Agents & ML: LLM agents, Skills, Function Calling, RAG, MCP, Evals, Context Engineering, PyTorch
Data & infra: PySpark, Palantir Foundry, AWS, Grafana, Prometheus, Dagster
Languages: Python, TypeScript, SQL
Contact

If the work is interesting, I'm around.

Applied AI, agent evaluation, forward deployed engineering, or something I haven't thought of yet.