Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions NEWS.md
Original file line number Diff line number Diff line change
@@ -1,2 +1,15 @@
# LibXLS.jl v1.0.0 Release Notes

* Complete rewrite as a native reader for legacy Excel xls files, wrapping
the libxls C library via the registered `libxls_jll` binaries (no
BinaryProvider, no build step).
* Full cell value support: numbers, strings, booleans, blank cells
(`missing`), error cells (new `CellError` type) and cached formula results.
* Date and time support: cell number formats (builtin ids and custom format
strings) determine whether a number is a `Date`, `DateTime` or `Time`,
honoring both the 1900 and the 1904 date system.
* Worksheet access via `getworksheet`/`wb[...]`, `size` and `ws[row, col]`.
* Tests use the test item framework; minimum supported Julia is 1.12.

# LibXLS.jl v0.0.1 Release Notes
* Initial release
93 changes: 85 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,20 +11,29 @@ binaries provided by `libxls_jll`, so it works without any Python or Java
dependency.

For modern xlsx files use [XLSX.jl](https://github.com/JuliaData/XLSX.jl)
instead; LibXLS deliberately only handles the legacy format.
instead; LibXLS deliberately only handles the legacy format. If you want one
package that reads both formats behind a single API, use
[ExcelReaders.jl](https://github.com/queryverse/ExcelReaders.jl), which is
built on LibXLS and XLSX.jl.

## Usage
## Installation

```julia
Pkg.add("LibXLS")
```

## Getting started

```julia
using LibXLS

wb = openxls("data.xls")

sheetnames(wb) # names of all sheets
sheetnames(wb) # names of all sheets
ws = getworksheet(wb, "Sheet1") # or by index: getworksheet(wb, 1)

nrows, ncols = size(ws)
ws[1, 1] # value of the cell in the first row and column
ws[1, 1] # value of the cell in the first row and column

close(wb)
```
Expand All @@ -37,7 +46,75 @@ openxls("data.xls") do wb
end
```

Cell values are returned as `Float64`, `String`, `Bool`, `DateTime`, `Time`,
`CellError` (for cells holding an Excel error such as `#DIV/0!`) or `missing`
(for blank cells). Whether a numeric cell holds a date is determined from the
cell's number format, honoring both the 1900 and the 1904 date system.
## API

### Workbooks

* `openxls(filepath)` opens an xls file and returns a `Workbook`. The file
format is validated from the file's content; opening an xlsx file gives an
error pointing to XLSX.jl. `openxls(f, filepath)` calls `f` on the workbook
and closes it afterwards.
* `close(wb)` releases the resources held by the C library; `isopen(wb)`
reports whether the workbook is still open. Workbooks also close themselves
when garbage collected.
* `sheetcount(wb)` — number of sheets, including hidden and empty ones.
* `sheetnames(wb)` — names of all sheets, in sheet order.
* `LibXLS.sheetname(wb, i)` / `LibXLS.sheetindex(wb, name)` — translate
between sheet indices and names.
* `LibXLS.isvisible(wb, index_or_name)` — whether a sheet is visible.
* `LibXLS.is1904(wb)` — whether the file uses the 1904 date system (cell
values already account for this).

### Worksheets

* `getworksheet(wb, index_or_name)` returns a `Worksheet`; `wb[1]` and
`wb["Sheet1"]` are shorthands. Worksheets are parsed on first access and
cached.
* `size(ws)` — dimensions of the used cell range as `(rows, columns)`.
* `ws[row, col]` — the value of a cell, using 1-based indices.
* `LibXLS.sheetname(ws)` / `LibXLS.sheetindex(ws)` — the sheet's name/index.

### Cell values

`ws[row, col]` returns plain Julia values:

| Excel cell | Julia value |
| --------------------------------------- | -------------------- |
| blank | `missing` |
| number | `Float64` |
| text | `String` |
| boolean | `Bool` |
| date (date-only number format) | `Dates.Date` |
| date + time | `Dates.DateTime` |
| time (time-only format, less than 24h) | `Dates.Time` |
| error (`#DIV/0!`, `#N/A`, ...) | `CellError` |

Formula cells return the cached result of the formula, mapped by the same
rules.

An xls file stores dates and times as plain numbers; whether a number denotes
a date is determined by the cell's number format, both for the builtin
formats and for custom format strings. Serial values convert correctly in
both the 1900 date system (including Excel's phantom 1900-02-29, which maps
to 1900-02-28) and the 1904 date system.

`CellError` wraps the BIFF error code of a cell holding an Excel error value
and prints as the corresponding error string:

| Code | Error |
| ------ | --------- |
| `0x00` | `#NULL!` |
| `0x07` | `#DIV/0!` |
| `0x0F` | `#VALUE!` |
| `0x17` | `#REF!` |
| `0x1D` | `#NAME?` |
| `0x24` | `#NUM!` |
| `0x2A` | `#N/A` |

## Limitations

* Reading only — for writing spreadsheets use
[XLSX.jl](https://github.com/JuliaData/XLSX.jl) (xlsx).
* Legacy xls only; xlsx files are rejected with a pointer to XLSX.jl.
* Defined names are not exposed by the underlying C library.
* Password-protected/encrypted workbooks are not supported.
2 changes: 1 addition & 1 deletion src/LibXLS.jl
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,8 @@ using libxls_jll: libxlsreader
export openxls, sheetcount, sheetnames, getworksheet, CellError

include("c.jl")
include("types.jl")
include("formats.jl")
include("types.jl")
include("workbook.jl")
include("worksheet.jl")

Expand Down
81 changes: 60 additions & 21 deletions src/formats.jl
Original file line number Diff line number Diff line change
@@ -1,18 +1,38 @@
# Detection of date/time number formats and conversion of Excel serial date
# values. A cell value in an xls file is just a Float64; whether it denotes a
# date or time is determined by the number format of the cell's XF record.
# date, a time or both is determined by the number format of the cell's XF
# record.

"""
CellFormatKind

Classification of a cell's number format: `FORMAT_NONE` for ordinary numbers,
`FORMAT_DATE` for date-only formats, `FORMAT_TIME` for time-only formats and
`FORMAT_DATETIME` for formats with both date and time components.
"""
@enum CellFormatKind begin
FORMAT_NONE
FORMAT_DATE
FORMAT_TIME
FORMAT_DATETIME
end

# Builtin number format ids that denote dates or times (same set xlrd uses):
# 14-22 date/time, 27-36 East Asian date, 45-47 elapsed time, 50-58 East
# Asian date variants.
const BUILTIN_DATE_FORMAT_IDS = Set{Int}([14:22; 27:36; 45:47; 50:58])
# 14-17 dates, 18-21 times, 22 date+time, 27-36 East Asian dates, 45-47
# elapsed times, 50-58 East Asian date variants.
const BUILTIN_FORMAT_KINDS = Dict{Int,CellFormatKind}(
Dict(i => FORMAT_DATE for i in [14:17; 27:36; 50:58])...,
Dict(i => FORMAT_TIME for i in [18:21; 45:47])...,
22 => FORMAT_DATETIME,
)

# Decide whether a custom number format string denotes a date/time. Quoted
# literals, escaped characters, padding/fill markers and bracket sections
# (colors like [Red], conditions like [<=100]) must be ignored; elapsed-time
# tokens like [h] or [ss] count as time. What remains is a date/time format
# iff it contains any of the date/time format codes y, m, d, h or s.
function is_date_format_string(fmt::AbstractString)
# Classify a custom number format string. Quoted literals, escaped characters,
# padding/fill markers and bracket sections (colors like [Red], conditions
# like [<=100]) must be ignored; elapsed-time tokens like [h] or [ss] count as
# time. In what remains, y and d denote a date component, h and s a time
# component, and m either of the two: minutes when next to an h or s code,
# months otherwise.
function format_string_kind(fmt::AbstractString)
stripped = IOBuffer()
i = firstindex(fmt)
n = lastindex(fmt)
Expand All @@ -36,27 +56,45 @@ function is_date_format_string(fmt::AbstractString)
i = nextind(fmt, i)
end
end
return occursin(r"[ymdhs]"i, String(take!(stripped)))
codes = String(take!(stripped))

# AM/PM markers denote a time component but their m is not a month.
hastime = occursin(r"AM/PM|A/P"i, codes)
codes = replace(codes, r"AM/PM|A/P"i => "")

hastime |= occursin(r"[hs]"i, codes)
hasdate = occursin(r"[yd]"i, codes)
# A run of m codes not adjacent to an h or s code means months, not
# minutes. The lookarounds exclude m itself so that only complete runs
# are considered.
hasdate |= occursin(r"(?<![hsm])m+(?![hsm])"i, replace(codes, r"[^a-zA-Z]" => ""))

hasdate && hastime && return FORMAT_DATETIME
hasdate && return FORMAT_DATE
hastime && return FORMAT_TIME
return FORMAT_NONE
end

function is_date_format(index::Integer, custom_formats::Dict{UInt16,String})
function format_kind(index::Integer, custom_formats::Dict{UInt16,String})
# A FORMAT record can redefine any index, including builtin ones, so the
# formats stored in the file take precedence over the builtin table.
haskey(custom_formats, index) && return is_date_format_string(custom_formats[index])
return Int(index) in BUILTIN_DATE_FORMAT_IDS
haskey(custom_formats, index) && return format_string_kind(custom_formats[index])
return get(BUILTIN_FORMAT_KINDS, Int(index), FORMAT_NONE)
end

const MILLISECONDS_PER_DAY = 86_400_000

"""
excel_serial_to_temporal(value, is1904)
excel_serial_to_temporal(value, kind, is1904)

Convert an Excel serial date/time value to a `DateTime`, or to a `Time` when
the value has no date component (`0 <= value < 1`). In the 1900 date system
the nonexistent date 1900-02-29 (serial 60) maps to 1900-02-28. Negative
values are not valid dates and are returned unchanged as `Float64`.
Convert an Excel serial date/time value to a `Date`, `DateTime` or `Time`,
depending on the format `kind` of the cell and the value: a date-only format
with no fractional part gives a `Date`, a time-only format with a value below
one day gives a `Time`, and everything else gives a `DateTime`. In the 1900
date system the nonexistent date 1900-02-29 (serial 60) maps to 1900-02-28.
Negative values are not valid dates and are returned unchanged as `Float64`.
"""
function excel_serial_to_temporal(value::Float64, is1904::Bool)
function excel_serial_to_temporal(value::Float64, kind::CellFormatKind, is1904::Bool)
value < 0 && return value
days = floor(Int, value)
ms = round(Int, (value - days) * MILLISECONDS_PER_DAY)
Expand All @@ -65,7 +103,7 @@ function excel_serial_to_temporal(value::Float64, is1904::Bool)
ms = 0
end
t = Time(0) + Millisecond(ms)
days == 0 && return t
days == 0 && kind != FORMAT_DATE && return t
if is1904
d = Date(1904, 1, 1) + Day(days)
elseif days == 60
Expand All @@ -75,5 +113,6 @@ function excel_serial_to_temporal(value::Float64, is1904::Bool)
else
d = Date(1899, 12, 30) + Day(days)
end
kind == FORMAT_DATE && ms == 0 && return d
return DateTime(d, t)
end
6 changes: 3 additions & 3 deletions src/types.jl
Original file line number Diff line number Diff line change
Expand Up @@ -50,10 +50,10 @@ mutable struct Workbook <: AbstractWorkbook
sheets_info::Vector{WorksheetInfo}
sheetname_index::Dict{String,Int}
sheets::Dict{Int,Worksheet}
xf_isdate::Vector{Bool} # per XF record: does its number format denote a date/time?
xf_kind::Vector{CellFormatKind} # per XF record: does its number format denote a date and/or time?

function Workbook(handle::Ptr{xlsWorkBook}, is1904::Bool, charset::String, sheets_info::Vector{WorksheetInfo}, sheetname_index::Dict{String,Int}, sheets::Dict{Int,Worksheet}, xf_isdate::Vector{Bool})
new_wb = new(handle, is1904, charset, sheets_info, sheetname_index, sheets, xf_isdate)
function Workbook(handle::Ptr{xlsWorkBook}, is1904::Bool, charset::String, sheets_info::Vector{WorksheetInfo}, sheetname_index::Dict{String,Int}, sheets::Dict{Int,Worksheet}, xf_kind::Vector{CellFormatKind})
new_wb = new(handle, is1904, charset, sheets_info, sheetname_index, sheets, xf_kind)
finalizer(close, new_wb)
return new_wb
end
Expand Down
60 changes: 57 additions & 3 deletions src/workbook.jl
Original file line number Diff line number Diff line change
Expand Up @@ -28,15 +28,15 @@ function Workbook(filepath::AbstractString)
end
end

xf_isdate = Vector{Bool}(undef, xlswb.xfs.count)
xf_kind = Vector{CellFormatKind}(undef, xlswb.xfs.count)
for i in 1:xlswb.xfs.count
xf_data = unsafe_load(xlswb.xfs.xf, i)
xf_isdate[i] = is_date_format(xf_data.format, custom_formats)
xf_kind[i] = format_kind(xf_data.format, custom_formats)
end

charset = xlswb.charset == C_NULL ? "" : unsafe_string(xlswb.charset)

return Workbook(handle, xlswb.is1904 != 0, charset, sheets_info, sheetname_index, Dict{Int,Worksheet}(), xf_isdate)
return Workbook(handle, xlswb.is1904 != 0, charset, sheets_info, sheetname_index, Dict{Int,Worksheet}(), xf_kind)
end

"""
Expand Down Expand Up @@ -79,6 +79,13 @@ function check_xls_file_format(filepath::AbstractString)
end
end

"""
close(wb::Workbook)

Close the workbook and release the resources held by the C library. Closing
an already closed workbook does nothing. Worksheets obtained from the
workbook must not be accessed afterwards.
"""
function Base.close(wb::Workbook)
if wb.handle != C_NULL
for ws in values(wb.sheets)
Expand All @@ -90,10 +97,34 @@ function Base.close(wb::Workbook)
return nothing
end

"""
isopen(wb::Workbook)

Whether the workbook has not been closed yet.
"""
Base.isopen(wb::Workbook) = wb.handle != C_NULL

"""
sheetcount(wb::Workbook)

The number of sheets in the workbook, including hidden and empty ones.
"""
sheetcount(wb::Workbook)::Int = length(wb.sheets_info)
"""
is1904(wb::Workbook)

Whether the workbook uses the 1904 date system (the default on classic Mac
versions of Excel) rather than the 1900 date system. Cell values already
account for this, so this is informational only.
"""
is1904(wb::Workbook)::Bool = wb.is1904
"""
sheetname(wb::Workbook, sheet_index)
sheetname(ws::Worksheet)

The name of the sheet with the given (1-based) index, or of the given
worksheet.
"""
sheetname(wb::Workbook, sheet_index::Integer)::String = wb.sheets_info[sheet_index].name

@inline is_valid_sheetindex(wb::Workbook, sheet_index::Integer) = 0 < sheet_index <= sheetcount(wb)
Expand All @@ -105,15 +136,38 @@ end
is_valid_sheetname(wb, sheet_name) || error("$sheet_name is not a valid sheet name.")
end

"""
sheetindex(wb::Workbook, sheet_name)
sheetindex(ws::Worksheet)

The (1-based) index of the sheet with the given name, or of the given
worksheet.
"""
@inline function sheetindex(wb::Workbook, sheet_name::AbstractString)::Int
check_valid_sheetname(wb, sheet_name)
return wb.sheetname_index[sheet_name]
end

"""
sheetnames(wb::Workbook)

The names of all sheets in the workbook, in sheet order.
"""
sheetnames(wb::Workbook)::Vector{String} = [sheetname(wb, i) for i in 1:sheetcount(wb)]
"""
isvisible(wb::Workbook, sheet_index_or_name)

Whether the sheet with the given index or name is visible (not hidden).
"""
isvisible(wb::Workbook, sheet_index::Integer)::Bool = wb.sheets_info[sheet_index].isvisible
isvisible(wb::Workbook, sheet_name::AbstractString)::Bool = isvisible(wb, sheetindex(wb, sheet_name))

"""
getworksheet(wb::Workbook, sheet_index_or_name) -> Worksheet

Return the worksheet with the given (1-based) index or name, parsing it on
first access. `wb[i]` and `wb["name"]` are shorthands for this function.
"""
function getworksheet(wb::Workbook, sheet_index::Integer)::Worksheet
wb.handle == C_NULL && error("Workbook is closed.")
if sheet_index ∉ keys(wb.sheets)
Expand Down
Loading
Loading