Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 14 additions & 2 deletions .github/workflows/nix.yml
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,9 @@ jobs:
- name: Run tests
run: nix develop --command cargo test --verbose

- name: Run tests (all features)
run: nix develop --command cargo test --all-features --verbose

no-std:
runs-on: ubuntu-latest
steps:
Expand All @@ -48,6 +51,9 @@ jobs:
- name: Build no_std + alloc (bare-metal target)
run: nix develop --command cargo build --no-default-features --features alloc --lib --target thumbv7em-none-eabi --verbose

- name: Build no_std + no_alloc + cp437g (bare-metal target)
run: nix develop --command cargo build --no-default-features --features cp437g --lib --target thumbv7em-none-eabi --verbose

fuzz:
runs-on: ubuntu-latest
steps:
Expand Down Expand Up @@ -108,15 +114,21 @@ jobs:
- name: Run clippy
run: nix develop --command cargo clippy --all-targets -- -D warnings

# The `#[cfg(feature = "alloc")]` gating means an unused import in a
# reduced feature set can only be caught by linting that set. Benches and
# The `#[cfg(feature = ...)]` gating means an unused import in a reduced
# feature set can only be caught by linting that set. Benches and
# integration targets are not feature-gated, so these are lib-scoped.
- name: Run clippy (no_std, no allocator)
run: nix develop --command cargo clippy --no-default-features --lib -- -D warnings

- name: Run clippy (no_std, alloc)
run: nix develop --command cargo clippy --no-default-features --features alloc --lib -- -D warnings

- name: Run clippy (no_std, cp437g)
run: nix develop --command cargo clippy --no-default-features --features cp437g --lib -- -D warnings

- name: Run clippy (all features)
run: nix develop --command cargo clippy --all-targets --all-features -- -D warnings

codegen:
runs-on: ubuntu-latest
steps:
Expand Down
3 changes: 3 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,9 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [2.1.0] - 2026-06-16
- Add an optional `cp437g` feature: the `CP437G` code page (CP437 overlaid with the IBM-Graphics glyphs at the C0 control byte range) for VGA text-mode rendering. Off by default.

## [2.0.1] - 2026-06-14
- Speed up the allocation-free `decode_byte` primitive (~2-4x) in `no_std`/no-allocator builds: the decode table now stores codepoints directly, so `decode_byte` lowers to a single indexed load. No API changes.

Expand Down
5 changes: 4 additions & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[package]
name = "yore"
version = "2.0.1"
version = "2.1.0"
authors = ["Andreas Liljeqvist <bonega@gmail.com>"]
edition = "2021"
categories = ["encoding"]
Expand All @@ -20,6 +20,9 @@ std = ["alloc"]
# `no_std` targets without an allocator; only the allocation-free char
# primitives (`encode_char`, `decode_byte`) remain available.
alloc = []
# Adds the CP437G code page: CP437 overlaid with the IBM-Graphics glyphs at the
# C0 control byte range, for VGA text-mode rendering. Off by default.
cp437g = []

[dependencies]

Expand Down
74 changes: 48 additions & 26 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

A Rust library for decoding and encoding character sets based on OEM code pages.

[![yore at crates.io](https://img.shields.io/badge/crates.io-2.0.1-blue)](https://crates.io/crates/yore)
[![yore at crates.io](https://img.shields.io/badge/crates.io-2.1.0-blue)](https://crates.io/crates/yore)
[![yore at docs.rs](https://docs.rs/yore/badge.svg)](https://docs.rs/yore)

# Features
Expand All @@ -11,50 +11,40 @@ A Rust library for decoding and encoding character sets based on OEM code pages.
* Easy-to-use API
* Broad range of [supported code pages](#supported-code-pages)
* Handles code pages with redefined ASCII characters (<0x80), such as '٪' in CP864
* `no_std` support, with or without an allocator, via cargo features
* Allocation-free `encode_char` / `decode_byte` primitives for embedded use
* [`no_std` support](#no_std), with or without an allocator — down to allocation-free `encode_char` / `decode_byte` primitives for embedded use

# Usage

Add `yore` to your `Cargo.toml` file.

```toml
[dependencies]
yore = "2.0.1"
yore = "2.1.0"
```

## `no_std`
## CP437G (VGA text-mode glyphs)

`yore` has three feature tiers, so it scales from std down to bare-metal targets
with no allocator:

| Cargo features | Environment | API |
| --- | --- | --- |
| `std` (default) | `std` | Full API + `std::error::Error` impls |
| `alloc` | `no_std` + allocator | Full API; `Error` impls omitted (`Display` stays) |
| *(none)* | `no_std`, no allocator | Allocation-free char primitives only |

The allocating `encode`/`decode` family returns owned `Cow` buffers and so
requires the `alloc` feature. Without it, only the allocation-free
`encode_char` / `decode_byte` primitives are available:
The optional `cp437g` feature adds the `CP437G` code page: CP437 overlaid with
the IBM-Graphics glyphs (☺ ♥ ♪ → ⌂ ...) at the C0 control byte range, as the VGA
BIOS glyph ROM renders them. It pairs well with `encode_char` for driving a text
buffer from a `no_std` kernel.

```toml
[dependencies]
# no_std with an allocator: keep the full Cow-returning API
yore = { version = "2.0.1", default-features = false, features = ["alloc"] }

# no_std without an allocator: char primitives only
yore = { version = "2.0.1", default-features = false }
yore = { version = "2.1.0", features = ["cp437g"] }
```

```rust
use yore::code_pages::CP850;
use yore::code_pages::CP437G;

// Encode/decode one character at a time, no allocation required.
assert_eq!(CP850.encode_char('A'), Some(b'A'));
assert_eq!(CP850.decode_byte(b'A'), 'A');
assert_eq!(CP437G.encode_char('☺'), Some(0x01));
assert_eq!(CP437G.decode_byte(0x01), '☺');
```

Some glyphs share a byte with an ASCII control character (`0x09` ○, `0x0A` ◙,
`0x0D` ♪). The ASCII fast-path still encodes `'\t'`/`'\n'`/`'\r'` to those
bytes, so intercept the source `char` first if you need newline semantics.

# Examples

## Using a specific code page
Expand Down Expand Up @@ -129,6 +119,38 @@ fn do_something(code_page: &dyn CodePage, bytes: &[u8]) {

Refer to the [bench crate](https://github.com/bonega/yore/blob/28198ff8d4e487a8f7e6a477fe7cbc19313618c0/benchmark/README.md) for more details.

# `no_std`

`yore` has three feature tiers, so it scales from std down to bare-metal targets
with no allocator:

| Cargo features | Environment | API |
| --- | --- | --- |
| `std` (default) | `std` | Full API + `std::error::Error` impls |
| `alloc` | `no_std` + allocator | Full API; `Error` impls omitted (`Display` stays) |
| *(none)* | `no_std`, no allocator | Allocation-free char primitives only |

The allocating `encode`/`decode` family returns owned `Cow` buffers and so
requires the `alloc` feature. Without it, only the allocation-free
`encode_char` / `decode_byte` primitives are available:

```toml
[dependencies]
# no_std with an allocator: keep the full Cow-returning API
yore = { version = "2.1.0", default-features = false, features = ["alloc"] }

# no_std without an allocator: char primitives only
yore = { version = "2.1.0", default-features = false }
```

```rust
use yore::code_pages::CP850;

// Encode/decode one character at a time, no allocation required.
assert_eq!(CP850.encode_char('A'), Some(b'A'));
assert_eq!(CP850.decode_byte(b'A'), 'A');
```

# Contributing

See [CONTRIBUTING.md](CONTRIBUTING.md) for development setup, benchmarking, and fuzzing.
49 changes: 49 additions & 0 deletions codegen/src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,55 @@ fn main() -> Result<()> {
generate_coder(&name, definition)?;
}

// CP437G: CP437 overlaid with the IBM-Graphics glyphs at the C0 control
// byte range, transcribed from unicode.org's IBMGRAPH.TXT.
//
// On VGA text mode the BIOS glyph ROM renders bytes 0x01-0x1F and 0x7F as
// pictographs (smiley, suits, arrows, ...) rather than control codes. This
// variant makes those glyphs representable so a kernel can drive the text
// buffer directly. 0x00 is left as NUL (the glyph ROM renders it blank).
// Some glyphs share a byte with an ASCII control (0x09 ○,
// 0x0A ◙, 0x0D ♪); callers that want to treat `\n` etc. as line-control
// should intercept the source `char` before calling `encode_char` (yore's
// ASCII fast-path still encodes `'\t'/'\n'/'\r'` to those same bytes).
let cp437g = {
let mut m = parsers::parse_unicode_dot_org("tables/unicode.org/CP437.txt")?;
m[0x01] = Some('\u{263A}'); // ☺
m[0x02] = Some('\u{263B}'); // ☻
m[0x03] = Some('\u{2665}'); // ♥
m[0x04] = Some('\u{2666}'); // ♦
m[0x05] = Some('\u{2663}'); // ♣
m[0x06] = Some('\u{2660}'); // ♠
m[0x07] = Some('\u{2022}'); // •
m[0x08] = Some('\u{25D8}'); // ◘
m[0x09] = Some('\u{25CB}'); // ○
m[0x0A] = Some('\u{25D9}'); // ◙
m[0x0B] = Some('\u{2642}'); // ♂
m[0x0C] = Some('\u{2640}'); // ♀
m[0x0D] = Some('\u{266A}'); // ♪
m[0x0E] = Some('\u{266B}'); // ♫
m[0x0F] = Some('\u{263C}'); // ☼
m[0x10] = Some('\u{25BA}'); // ►
m[0x11] = Some('\u{25C4}'); // ◄
m[0x12] = Some('\u{2195}'); // ↕
m[0x13] = Some('\u{203C}'); // ‼
m[0x14] = Some('\u{00B6}'); // ¶
m[0x15] = Some('\u{00A7}'); // §
m[0x16] = Some('\u{25AC}'); // ▬
m[0x17] = Some('\u{21A8}'); // ↨
m[0x18] = Some('\u{2191}'); // ↑
m[0x19] = Some('\u{2193}'); // ↓
m[0x1A] = Some('\u{2192}'); // →
m[0x1B] = Some('\u{2190}'); // ←
m[0x1C] = Some('\u{221F}'); // ∟
m[0x1D] = Some('\u{2194}'); // ↔
m[0x1E] = Some('\u{25B2}'); // ▲
m[0x1F] = Some('\u{25BC}'); // ▼
m[0x7F] = Some('\u{2302}'); // ⌂
m
};
generate_coder("CP437G", cp437g)?;

let whatwg_encodings = [874, 1250, 1251, 1252, 1253, 1254, 1255, 1256, 1257, 1258];
for cp in whatwg_encodings {
let name = format!("CP{}", cp);
Expand Down
Loading
Loading