Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
87 changes: 87 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,93 @@

## [Unreleased]

## [0.37.0] - 2026-07-25

### Added

- **`Lstm` layer.** Single-layer, batch-first LSTM mirroring `Gru`'s unroll-at-trace-time design and
built from existing primitives only (matmul/narrow/sigmoid/tanh/multiply) — no new `TensorOps` op,
and it traces to StableHLO with no dedicated converter. Gate order `i,f,g,o` and dual biases match
`torch.nn.LSTM`. Adds `LstmState(h, c)` with an explicit caller-owned `step(xt, state, ctx)` API —
required by transducer prediction networks (RNN-T/TDT) and exactly the shape that lowers to a
fixed-shape single-step StableHLO graph — plus an `initialState(batch, ctx, dtype)` helper. (PR #824)
- **Learning-rate schedules and mutable optimizer `lr`.** `AdamOptimizer` and `SgdOptimizer` expose
`lr` as a public `var`, so training loops can adjust it between steps without recreating the
optimizer and losing the moment estimates. A stateless `LrSchedule` fun interface plus a
`linearWarmupCosineDecay` factory cover the GPT-style pretraining recipe: linear ramp from
`initialLr` to `peakLr` over the warmup steps, then cosine decay to `minLr`. (PR #866, issue #865)
- **BREAKING (Java callers): optional bias in `Linear`.** `initBias` is now nullable with a `null`
default, creating a bias-less projection (`y = x W^T`) — the equivalent of PyTorch's
`nn.Linear(bias=False)`, needed for GPT-style `qkv_bias=false` projections and weight-tied output
heads. A bias-less layer registers only its weight parameter, so parameter counts and checkpoints
match architectures defined without bias. Adds a `biasOrNull()` helper next to the throwing
`bias()` accessor. **The `@JvmOverloads` overload set changes:** the 4-arg
`(in, out, weights, bias)` convenience overload is replaced by `(in, out, weights)`, so Java code
binding the 4-arg overload fails to compile (and breaks binary compatibility) until it moves to the
full constructor. Kotlin callers bind the full constructor and are unaffected. (PR #870)
- **`Linear` is `open`.** Adapter-style layers (LoRA and similar) can subclass it and augment
`onForward` or `params` while reusing the base projection. No behavior change — the override points
were already open via `Module`. (PR #875)
- **`androidNative` targets for the IO modules.** `skainet-io-core` gains `androidNativeArm64` and
`androidNativeArm32` (via a 64-bit-only `native64Main` source set), and `skainet-io-safetensors`
gains the `androidNative` target set; the latter also drops unused `compile-core`/`compile-dag`
dependencies. (PRs #836, #842, #845)
- **OpenSSF Scorecard workflow and badge.** (PR #814)

### Changed

- **`Dropout` now actually drops.** It was an identity placeholder in both phases; it now applies
inverted dropout under a training-phase context — each element is zeroed with probability `p` and
survivors are scaled by `1/(1-p)`, so the expected activation is preserved and inference needs no
rescaling. The mask is a constant tensor combined with an element-wise multiply, so gradients flow
to surviving inputs without a dedicated autograd rule; an injectable `Random` enables reproducible
masks. Models that relied on `Dropout` being a no-op will see different (correct) training-time
activations. (PR #867, issue #861)
- **Toolchain and dependency bumps.** Kotlin 2.4.10 (JVM plugin and `kotlinx.serialization` plugin
aligned), AGP 9.3.1, Shadow 9.6.1, Kover 0.9.9, JUnit Jupiter 6.1.2, and Logback 1.6.0. No public
API or behavior changes. (PRs #818, #819, #820, #826, #829, #840, #856, #873, #874, #811)
- **Supply-chain hardening of CI and the docs image.** All GitHub Actions across every workflow are
now pinned to commit hashes, and the documentation `Dockerfile` pins its base image by digest and
its npm packages (Antora CLI, site-generator, lunr-extension, mermaid-cli, asciidoctor-kroki) to
exact versions from the npm registry, making documentation builds reproducible. (PRs #816, #821,
#827, #830, #831, #832, #834, #837, #838, #846, #848, #868)
- **CI `allTests` split into parallel per-target jobs**, ending the recurring out-of-memory flakes on
the combined job. (PR #854)
- **Docs deployment publishes on release**, and fork PRs no longer fail on the comment step. (PR #850)

### Fixed

- **`scaledDotProductAttention` silently discarded the attention pattern at its default scale.** The
op's `scale` parameter defaults to `0f`, documented as meaning `1/sqrt(headDim)`, but the CPU
backend applied it literally — every score was multiplied by zero, so the softmax collapsed to a
uniform average. The CPU kernel now resolves `scale == 0f` to `1/sqrt(headDim)`. Any model calling
SDPA without an explicit scale was producing attention-free output. (PR #880, issue #860)
- **Three autograd bugs that silently froze or crashed training.** (PR #877, issues #862, #863, #864)
- `CrossEntropyLoss`: the index-target path built its result host-side via
`tensorDataFactory` + `fromData`, detaching the tape so no gradient reached the predictions, and
the soft-target path only recorded when targets lived on the recording context. Both now compute
the NLL with differentiable ops dispatched through the predictions' ops (a constant one-hot for
the index path), keeping the tape connected.
- `softmax`/`logSoftmax` backward: a negative `dim` (e.g. `softmax(dim = -1)`) was passed straight
to `broadcastToInput`, which reserves `-1` as a no-unsqueeze sentinel, so the reduced axis was
never re-expanded and backward crashed for rank ≥ 3. The dim is now normalized.
- `varianceBackward`: the reduced mean and upstream kept the reduced shape, so the
subtract/multiply could not broadcast for rank ≥ 2, and `N` used the total volume instead of the
axis size. Both are now expanded over the reduced axis and `N` is the axis size.
- **`argMax` traced through the `dag {}` builder emitted invalid StableHLO.**
`GraphDsl.inferDagOutputSpecs` had no `argMax` case, so it echoed operand-0's shape and dtype; the
converter then faithfully emitted an integer literal for an `f32` tensor (rejected by
`iree-compile`) and a reduce whose result kept the reduced dim. The builder now infers a reduced
shape with an `Int32` index dtype, matching the `VoidTensorOps` path that was already correct —
which is why only the raw `dag {}` path broke. (PR #878, issue #876)
- **Legacy `tokenizer.json` files without `model.type` failed to load.** Files such as
`openai-community/gpt2` omit the field; `TokenizerFactory.fromTokenizerJson` now infers the type
from structure — a `merges` list is unique to BPE — and routes them to
`QwenByteLevelBpeTokenizer` instead of throwing. (PR #879, issue #858)
- **CPU `gather` threw on multi-dimensional indices.** An `[N, L]` index tensor needs one coordinate
per dimension for a flat `data[i]` access; indices are now read in row-major order via the
contiguous buffer, falling back to unravel. (PR #879, issue #859)

## [0.36.0] - 2026-07-11

### Added
Expand Down
26 changes: 17 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ Add the core dependencies (Gradle Kotlin DSL):
```kotlin
dependencies {
// Recommended: import the umbrella BOM and drop versions on the engine modules.
implementation(platform("sk.ainet:skainet-bom:0.36.0"))
implementation(platform("sk.ainet:skainet-bom:0.37.0"))

implementation("sk.ainet.core:skainet-lang-core")
implementation("sk.ainet.core:skainet-backend-cpu")
Expand Down Expand Up @@ -287,17 +287,20 @@ val withoutLabel = dataPipeline<RawDataset>()

---

## What's New in 0.36.0
## What's New in 0.37.0

- **Kotlin 2.4.0 toolchain** — the framework now builds on Kotlin 2.4.0, with KSP 2.3.10 and Dokka 2.2.0 aligned to the new compiler. No public API changes.
- **`permute` replay fix (`ComputeGraphExecutor`)** — a traced `permute(t, axes)` now replays with its recorded axes instead of being dispatched as a plain last-two-dims transpose. Rank-3+ permutations (e.g. multi-head attention's heads/sequence swap in full-sequence encoder/prefill traces) previously produced the wrong layout; decode paths were unaffected, which is why the bug hid.
- **REUSE / SPDX license-compliance setup** — `REUSE.toml` + `LICENSES/`, a CI compliance workflow, and a REUSE status badge, so the repository is machine-verifiable against the [REUSE](https://reuse.software/) specification.
- **`skainet-data` POM coordinate alignment** — data-module POM coordinates and display names now match their module names (see CHANGELOG for the `skainet-data-simple` artifactId note).
- **`Lstm` layer** — single-layer, batch-first LSTM built from existing primitives only (no new `TensorOps` op, traces to StableHLO without a dedicated converter), with `torch.nn.LSTM`-compatible gate order and an explicit caller-owned `LstmState` + `step()` API for transducer prediction networks.
- **Training essentials** — `Dropout` now performs real inverted dropout under a training-phase context (it was an identity placeholder), optimizers expose a mutable `lr` plus a `linearWarmupCosineDecay` LR schedule, `Linear` supports bias-less projections (`nn.Linear(bias=False)` equivalent) and is `open` for LoRA-style adapters.
- **Attention scale fix** — `scaledDotProductAttention` at its default scale multiplied every score by zero on the CPU backend, collapsing softmax to a uniform average; it now resolves to `1/sqrt(headDim)` as documented.
- **Autograd correctness** — `CrossEntropyLoss` no longer detaches the tape (gradients reached the predictions in neither target path), and `softmax`/`logSoftmax`/`variance` backward now work for rank ≥ 3.
- **Android native IO** — `skainet-io-core` and `skainet-io-safetensors` gain `androidNative` targets (arm64 and arm32).
- **Reproducible, hardened CI** — every GitHub Action pinned to a commit hash, the docs Docker image pinned by digest with exact npm package versions, and `allTests` split into parallel per-target jobs to end OOM flakes.

### Previously, in 0.35.0
### Previously, in 0.36.0

- **`argMax(dim)` tensor op** — index of the maximum along a dimension (ties → lowest index), lowered to StableHLO as a single op (no new primitive) plus an eager CPU kernel.
- **URI-backed data sources** — `skainet-data-source` module: `file://`, `https://`, and Hugging Face URIs, raw-format parsers (CSV/TSV/JSON/JSONL), suspendable data pipelines.
- **Kotlin 2.4.0 toolchain** — KSP 2.3.10 and Dokka 2.2.0 aligned to the new compiler. No public API changes.
- **`permute` replay fix (`ComputeGraphExecutor`)** — a traced `permute(t, axes)` now replays with its recorded axes instead of being dispatched as a plain last-two-dims transpose.
- **REUSE / SPDX license-compliance setup** — `REUSE.toml` + `LICENSES/`, a CI compliance workflow, and a REUSE status badge.

See [CHANGELOG.md](CHANGELOG.md) for details and the full release history.

Expand All @@ -323,6 +326,11 @@ We love contributions! Whether it's a new operator, documentation, or a bug fix:

Browse the full codebase documentation on [DeepWiki](https://deepwiki.com/SKaiNET-developers/SKaiNET).

### Contributors (0.37.0)

- **Michal Harakal** ([@michalharakal](https://github.com/michalharakal)) — `Lstm` layer (#824), `Dropout` masking (#867), LR schedules (#866), optional/open `Linear` (#870, #875), SDPA scale fix (#880), autograd fixes (#877), `argMax` DAG spec (#878), tokenizer + `gather` fixes (#879), Android native IO targets (#836, #842, #845)
- **[@MacOS](https://github.com/MacOS)** — OpenSSF Scorecard workflow and badge (#814), commit-hash pinning across all CI workflows and reproducible docs Docker image (#816, #821, #827, #830–#838, #846, #848, #868)

### Contributors (0.36.0)

- **[@MacOS](https://github.com/MacOS)** — REUSE compliance CI workflow and status badge (#806, #807)
Expand Down
2 changes: 1 addition & 1 deletion docs/antora.yml
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ asciidoc:
framework_name: SKaiNET
# Current SKaiNET release — bump once per release; referenced as
# {skainet_version} in dependency snippets (blocks need subs="attributes+").
skainet_version: 0.36.0
skainet_version: 0.37.0
ksp_version: 2.2.21-2.0.5
dokka_version: 2.1.0
asciidoctorj_version: 3.0.0
2 changes: 1 addition & 1 deletion gradle.properties
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
GROUP=sk.ainet.core
VERSION_NAME=0.36.0
VERSION_NAME=0.37.0
POM_DESCRIPTION=SKaiNET

POM_URL=https://github.com/SKaiNET-developers/skainet/
Expand Down
Loading