# Building an iOS SDK for Bluetooth Low Energy hardware

> iOS SDK lead on a Bluetooth Low Energy device SDK: the connection layer, the device messaging protocol, firmware updates over the air, and the public API.

- Project: Connected Device SDK
- Role: iOS SDK lead
- Period: 2017–2020
- Topics: Swift, CoreBluetooth, BLE, device protocols, firmware updates, SDK API design
- Note: The client and the product are not named.

URL: https://jovev.com/work/connected-device-sdk
Author: Dragan Jovev (https://jovev.com)
Updated: 2026-09-30

## Context

A consumer hardware company was building a line of Bluetooth Low Energy devices and needed mobile SDKs so that applications could talk to them. The SDK sat between an application and the hardware: it discovered devices, connected to them, exchanged messages with the firmware, and updated that firmware over the air.

There were two audiences, in that order. First the company's own companion applications. Then developers outside the company, building on the same hardware. The second audience never arrived, but it shaped the API from the first day.

I was part of the cross-platform engineering group from 2017 to 2020, close to two years of it on the iOS SDK. iOS was one workstream; a separate team built the Android SDK, and a third wrote the device firmware — more than twenty engineers across the three.

## Role

I led the iOS SDK. I owned the Bluetooth connection layer, the messaging protocol the SDK and the firmware spoke, and over-the-air firmware updates. I also owned the shape of the public API on iOS — what an integrator imports and calls.

The API design direction was mine, and the vocabulary the two platform SDKs shared was agreed against it. Getting it adopted meant convincing the product owner as much as the other engineering teams.

## Constraints

- **The link is slow and small.** Bluetooth Low Energy is built for short, infrequent messages, around twenty bytes of payload per packet. A firmware image is neither short nor infrequent.
- **The hardware was not ours.** The firmware belonged to another team on its own schedule, and real devices were scarce. The thing the SDK talked to changed underneath it.
- **Two audiences with different tolerances.** An internal team can be told about a breaking change in a meeting. External developers find out when their build stops compiling.
- **People do not sit still.** The phone and the device are both carried. Anything that takes minutes has to survive someone walking out of range.

## Architecture

```
Application
   │
   ▼
Public API      scanner ─► token ─► device controller ─► firmware controller
   │
   ▼
Message layer   topics, publish/subscribe, request/response
   │
   ▼
Connection      CoreBluetooth: scan, connect, discover, read, write
   │
   ▼
  BLE ─ ─ ─ ─►  device firmware
```

### Discovery and identity

Scanning is asynchronous and reports results as they arrive. What it produces is not a connection but a token: a stable identifier and a product code. The token is the only thing that crosses from discovery into connection, so an application holds an identifier it can store and compare, and platform objects stay out of the public API.

### The connection layer

Connecting is not one step. The SDK connects to the device, discovers its services, then discovers the characteristics of each service. Only then does it report success to the caller, with a timeout covering the whole sequence rather than the first step alone.

Several reads can be outstanding at once, so each response is matched back to the caller that asked for it, rather than delivered to whoever happens to be listening.

### The message layer

Above the connection sits a topic-based protocol. A topic is a path of identifiers; a subscriber registers a pattern, with wildcards for one level or many. Three interaction shapes run over the same connection: publish and subscribe for streams the device initiates, request and response with correlation for commands the application initiates, and listing so an application can ask a device what it offers. Each message carries a priority and a quality-of-service level, so a command does not queue behind a stream of telemetry.

## Key decisions

### Shared concepts, platform-native APIs

The iOS and Android SDKs agreed on the vocabulary — scanner, token, controller, firmware controller — and on nothing else. Each was free to be idiomatic on its own platform. Documentation, design conversations and mental models transferred between the teams; API shapes did not. The alternative was one API shape ported to both, which would have made at least one of the two SDKs feel foreign to the developers it was written for.

### Keeping the public surface small

Everything public is a promise. Internal integrators tolerate churn; external ones do not, and the plan was always to open the SDK up eventually. So the public surface stayed deliberately narrow — a scanner, a token, a device controller, a firmware controller — and everything else stayed internal, where it could be rewritten without asking anyone's permission.

### Developer experience

- **Documented API.** Every public type and method carried documentation, written as part of the change rather than collected afterwards.
- **An error model, not error strings.** Failures carried a category, and categories were split by layer — device, application, cloud — so an integrator could tell whose problem a failure was before reading the message.
- **A reference app.** A working integration that served two purposes: it proved the API was usable before anyone else had to use it, and it was the thing you hand an integrator to copy from.

## Tradeoffs

Reporting a connection as ready only after full service and characteristic discovery makes connecting measurably slower, and the caller can do nothing while it runs. Reporting the link as soon as it is established would have felt faster. It would also have created a state where the application believes it is connected and every call fails. We took the slower connect: "connected" means usable.

## Failure modes

### A firmware image against a twenty-byte packet

A firmware image is a few megabytes. A packet carries around twenty bytes. An update is therefore tens of thousands of packets, and a full transfer takes about two minutes. Two things follow. Progress reported in bytes is not a nicety; it is the only evidence the user has that the update has not died. And any design upstream that quietly assumes the link is fast has to be found and redone, because the physical limit is not negotiable.

### Updates that run while life happens

The failure that mattered most was not a bug. It was the assumption that the phone stays near the device. An update takes minutes, and in those minutes someone puts the phone down, walks to another room, or leaves the building. A transfer has to be able to fail in the middle and be resumed or restarted without leaving the device in a state nobody can reach.

That class of problem does not appear in a unit test, and it does not appear on a bench with the device sitting next to the laptop. It appears when a person carries the phone down a corridor.

## Outcome and lessons

Dozens of applications shipped on the SDK, among them the one that shipped alongside the hardware itself. The public developer release never happened: the company wound down the product line, and the hardware went with it.

What I would do differently:

- **Treat the public API as the expensive decision.** Implementation behind a good API can be replaced at any time. The API itself cannot, once anything depends on it. It deserved more design time than the code behind it, and it got less.
- **Plan for hardware you don't control.** Another team's firmware, another team's schedule, and not enough real devices. I treated that as friction to work around. It was a constraint to design for.
- **Test on real hardware, in real conditions, early.** The failures that mattered only showed up when someone carried the phone out of range in the middle of a transfer. A device on the desk beside you will never show you those.
