---
title: "hackeps 2025: how we came second in the eurecat challenge"
description: "at hackeps 2025 we built a cluster platform on aws and google cloud: why we scrapped the go halfway through the hackathon and how what we ended up with works."
date: 2026-10-08
author: "joel taylor pedrós"
lang: en
url: https://joeltaylor.business/blog/hackeps-2025-eurecat-challenge
translations:
  ca: https://joeltaylor.business/blog/hackeps-2025
  es: https://joeltaylor.business/blog/hackeps-2025-reto-eurecat
image: https://joeltaylor.business/_next/static/media/cover.1iurd3q0ku09g.webp
---

# hackeps 2025: how we came second in the eurecat challenge

![the app's dashboard with three active clusters: one docker swarm cluster with three nodes and two k3s clusters with two nodes each.](https://joeltaylor.business/_next/static/media/cover.1iurd3q0ku09g.webp)

hackeps 2025 took place on 22 and 23 november, and i took part with maria aliet. it's the hackathon of the escola politècnica superior at the university of lleida: ninth edition, more than 230 sign-ups and 24 hours. we chose the eurecat challenge, which asked for a platform to deploy and manage hybrid computing clusters, with machines on aws and google cloud plus edge devices, all from one place.

we came second. but halfway through the hackathon we had nothing that worked, and we deleted everything to start again.

## the plan with go, connect and protocol buffers

we started as if we had a month. a go backend, with connect-go and protocol buffers, so the types would be shared end to end between the server and the client.

on paper it makes sense. in practice, every new endpoint needs a definition in the `.proto` file, the go code and the client code generated, and the handler implemented, before any button does anything. at the halfway mark we had a beautiful architecture diagram and zero finished features.

we did a hard reset and moved everything to t3:

- next.js and trpc, for types shared between client and server without generating code.
- postgresql with prisma, so we could change the schema without headaches.
- tailwind and shadcn, for a decent interface in minutes.

nothing of the go is left in the repository, which starts straight with the new version. what we presented is 3,832 lines of our own typescript in 60 files, 1,494 of them on the server, plus 5,745 lines of shadcn components we didn't write.

the cut also shows in the database schema. the first one had four possible orchestrators: k3s, nomad, docker swarm and kubernetes. a later migration removed nomad and kubernetes, since there wouldn't be time to build them.

## how a cluster is created on aws and google cloud

first you save each provider's credentials: the aws access key and secret, or the json for a google cloud service account. then you create a cluster. you pick docker swarm or k3s, add nodes with a provider and an instance type, and mark one as the master.

on creation, a trpc mutation saves the cluster and its nodes in the pending state inside a transaction, and calls `provisionCluster`, which does the work in three stages.

![five-step diagram: create the machines, wait 15 seconds, ask for the public ip, try ssh every 6 seconds for up to 50 attempts, and install swarm or k3s. the fourth step is highlighted. below, the node's state goes from pending to provisioning, and ends at active or failed.](https://joeltaylor.business/_next/static/media/aprovisionament.en.2yftzswz1grzy.webp)

_each node's state is saved in the database, and the dashboard shows it when you refresh._

first it launches the machines. the server generates a 4,096-bit rsa key pair per cluster. on aws, the public key goes in through the user data, a script that adds it to `authorized_keys` when the machine boots. on google cloud, it goes in through the instance's `ssh-keys` metadata. then it waits 15 seconds and asks each provider for the public ip. finally it connects over ssh, installs the orchestrator and marks the node as active.

asking google cloud for the public ip was one of the last things we added, and until then no google cloud node got past the provisioning state.

node.js generates rsa keys, but it can't export them in the openssh format that aws and google cloud expect. so there's a hand-written function that reads the pem key byte by byte, pulls out the modulus and the exponent and assembles the `ssh-rsa` key as rfc 4253 describes it.

the private key, on the other hand, ends up in the database in plain text, next to a comment that says "in a real app, this field should be encrypted :)".

## polling over ssh

we didn't have time to set up a task queue, so we improvised. a loop tries to connect over ssh to each new machine and runs an `echo` on it. if that fails, it waits six seconds and tries again. the node only becomes active once it answers. without the logging, it's this:

```ts
async function waitForSSH(ip: string, user: string, privateKey: string, maxRetries = 50) {
  for (let i = 0; i < maxRetries; i++) {
    try {
      await executeRemoteCommand(ip, user, privateKey, ["echo 'SSH Ready'"]);
      return;
    } catch (e) {
      await new Promise((r) => setTimeout(r, 6000));
    }
  }
  throw new Error(`SSH Connection timed out after ${maxRetries} attempts`);
}
```

the limit was 20 attempts, about two minutes, and we had to raise it to 50, about five.

all of this happens inside the same http request that creates the cluster, which doesn't respond until the last node is done. so you're not left staring at the form, the button fires the mutation and takes you to the dashboard without waiting for the response.

one of the commits is called "parallelisation of provisioning", but the code is still sequential. what changed is the order. before, each node was created and then waited 15 seconds for its ip. after, they're all launched, there's a single 15-second wait and all the ips are collected. a comment in the code itself admits it: "for safety in a hackathon (rate limits), let's keep it sequential but fast".

deploying an application also goes over ssh, always to the master. with swarm, the server runs `docker service create` there with the image you give it. with k3s, it encodes the manifest in base64, decodes it on the machine and runs `kubectl apply`. that way the yaml arrives intact, without fighting the quotes inside a shell command.

## what was left half done

the code also shows what never got working:

- with swarm, each node runs `docker swarm init` on its own, and with k3s it only gets installed on the master. the nodes don't join each other, because none of them runs `join`.
- edge devices are registered with an ip and a user, and marked as active if they have an ip, without anything connecting to them.
- on aws there's a single ami and a single security group hardcoded. the ami is the us-west-2 one and gets used in any region, even though each ami only exists in the region where it was created.

the readme, written two days later, also lists what didn't make it. the ai part of the challenge, picking a provider and region from a natural-language request, was designed but never wired up. and the first improvement on the list is a queue with redis and bullmq to take provisioning out of the http request. it's the piece the ssh polling had to stand in for.
