From 8c252285e8c84b25e39ed05d12117b5a32a56aa0 Mon Sep 17 00:00:00 2001 From: David Anthoff Date: Fri, 28 Aug 2026 19:36:25 -0700 Subject: [PATCH] Pre-registration polish: Date values, API docs, layout and coverage tests - Cells with a date-only number format now come back as Date rather than DateTime (matching XLSX.jl); time-only formats keep returning Time for values below one day. Format classification distinguishes date, time and datetime formats for builtin ids and custom format strings. - Add a C struct layout test item asserting the size of every mirrored libxls struct on 64-bit platforms, plus a check that the loaded C library is the 1.6.x series (idea from PR #15). - Document the full API in the README (workbook/worksheet functions, cell value mapping, date systems, error codes, limitations) and add docstrings to all public functions; write v1.0.0 release notes. - Add tests for the C error reporting path, corrupt-file handling, serial rounding carry and other previously uncovered branches. Co-Authored-By: Claude Fable 5 --- NEWS.md | 13 ++++ README.md | 93 +++++++++++++++++++++++++--- src/LibXLS.jl | 2 +- src/formats.jl | 81 ++++++++++++++++++------- src/types.jl | 6 +- src/workbook.jl | 60 ++++++++++++++++++- src/worksheet.jl | 12 +++- test/test_libxls.jl | 143 ++++++++++++++++++++++++++++++++++++-------- 8 files changed, 348 insertions(+), 62 deletions(-) diff --git a/NEWS.md b/NEWS.md index 4b54864..022b80f 100644 --- a/NEWS.md +++ b/NEWS.md @@ -1,2 +1,15 @@ +# LibXLS.jl v1.0.0 Release Notes + +* Complete rewrite as a native reader for legacy Excel xls files, wrapping + the libxls C library via the registered `libxls_jll` binaries (no + BinaryProvider, no build step). +* Full cell value support: numbers, strings, booleans, blank cells + (`missing`), error cells (new `CellError` type) and cached formula results. +* Date and time support: cell number formats (builtin ids and custom format + strings) determine whether a number is a `Date`, `DateTime` or `Time`, + honoring both the 1900 and the 1904 date system. +* Worksheet access via `getworksheet`/`wb[...]`, `size` and `ws[row, col]`. +* Tests use the test item framework; minimum supported Julia is 1.12. + # LibXLS.jl v0.0.1 Release Notes * Initial release diff --git a/README.md b/README.md index 35fc316..da068fc 100644 --- a/README.md +++ b/README.md @@ -11,20 +11,29 @@ binaries provided by `libxls_jll`, so it works without any Python or Java dependency. For modern xlsx files use [XLSX.jl](https://github.com/JuliaData/XLSX.jl) -instead; LibXLS deliberately only handles the legacy format. +instead; LibXLS deliberately only handles the legacy format. If you want one +package that reads both formats behind a single API, use +[ExcelReaders.jl](https://github.com/queryverse/ExcelReaders.jl), which is +built on LibXLS and XLSX.jl. -## Usage +## Installation + +```julia +Pkg.add("LibXLS") +``` + +## Getting started ```julia using LibXLS wb = openxls("data.xls") -sheetnames(wb) # names of all sheets +sheetnames(wb) # names of all sheets ws = getworksheet(wb, "Sheet1") # or by index: getworksheet(wb, 1) nrows, ncols = size(ws) -ws[1, 1] # value of the cell in the first row and column +ws[1, 1] # value of the cell in the first row and column close(wb) ``` @@ -37,7 +46,75 @@ openxls("data.xls") do wb end ``` -Cell values are returned as `Float64`, `String`, `Bool`, `DateTime`, `Time`, -`CellError` (for cells holding an Excel error such as `#DIV/0!`) or `missing` -(for blank cells). Whether a numeric cell holds a date is determined from the -cell's number format, honoring both the 1900 and the 1904 date system. +## API + +### Workbooks + +* `openxls(filepath)` opens an xls file and returns a `Workbook`. The file + format is validated from the file's content; opening an xlsx file gives an + error pointing to XLSX.jl. `openxls(f, filepath)` calls `f` on the workbook + and closes it afterwards. +* `close(wb)` releases the resources held by the C library; `isopen(wb)` + reports whether the workbook is still open. Workbooks also close themselves + when garbage collected. +* `sheetcount(wb)` — number of sheets, including hidden and empty ones. +* `sheetnames(wb)` — names of all sheets, in sheet order. +* `LibXLS.sheetname(wb, i)` / `LibXLS.sheetindex(wb, name)` — translate + between sheet indices and names. +* `LibXLS.isvisible(wb, index_or_name)` — whether a sheet is visible. +* `LibXLS.is1904(wb)` — whether the file uses the 1904 date system (cell + values already account for this). + +### Worksheets + +* `getworksheet(wb, index_or_name)` returns a `Worksheet`; `wb[1]` and + `wb["Sheet1"]` are shorthands. Worksheets are parsed on first access and + cached. +* `size(ws)` — dimensions of the used cell range as `(rows, columns)`. +* `ws[row, col]` — the value of a cell, using 1-based indices. +* `LibXLS.sheetname(ws)` / `LibXLS.sheetindex(ws)` — the sheet's name/index. + +### Cell values + +`ws[row, col]` returns plain Julia values: + +| Excel cell | Julia value | +| --------------------------------------- | -------------------- | +| blank | `missing` | +| number | `Float64` | +| text | `String` | +| boolean | `Bool` | +| date (date-only number format) | `Dates.Date` | +| date + time | `Dates.DateTime` | +| time (time-only format, less than 24h) | `Dates.Time` | +| error (`#DIV/0!`, `#N/A`, ...) | `CellError` | + +Formula cells return the cached result of the formula, mapped by the same +rules. + +An xls file stores dates and times as plain numbers; whether a number denotes +a date is determined by the cell's number format, both for the builtin +formats and for custom format strings. Serial values convert correctly in +both the 1900 date system (including Excel's phantom 1900-02-29, which maps +to 1900-02-28) and the 1904 date system. + +`CellError` wraps the BIFF error code of a cell holding an Excel error value +and prints as the corresponding error string: + +| Code | Error | +| ------ | --------- | +| `0x00` | `#NULL!` | +| `0x07` | `#DIV/0!` | +| `0x0F` | `#VALUE!` | +| `0x17` | `#REF!` | +| `0x1D` | `#NAME?` | +| `0x24` | `#NUM!` | +| `0x2A` | `#N/A` | + +## Limitations + +* Reading only — for writing spreadsheets use + [XLSX.jl](https://github.com/JuliaData/XLSX.jl) (xlsx). +* Legacy xls only; xlsx files are rejected with a pointer to XLSX.jl. +* Defined names are not exposed by the underlying C library. +* Password-protected/encrypted workbooks are not supported. diff --git a/src/LibXLS.jl b/src/LibXLS.jl index 5eaeb46..f33c9b0 100644 --- a/src/LibXLS.jl +++ b/src/LibXLS.jl @@ -6,8 +6,8 @@ using libxls_jll: libxlsreader export openxls, sheetcount, sheetnames, getworksheet, CellError include("c.jl") -include("types.jl") include("formats.jl") +include("types.jl") include("workbook.jl") include("worksheet.jl") diff --git a/src/formats.jl b/src/formats.jl index 4927ab8..90f2e97 100644 --- a/src/formats.jl +++ b/src/formats.jl @@ -1,18 +1,38 @@ # Detection of date/time number formats and conversion of Excel serial date # values. A cell value in an xls file is just a Float64; whether it denotes a -# date or time is determined by the number format of the cell's XF record. +# date, a time or both is determined by the number format of the cell's XF +# record. + +""" + CellFormatKind + +Classification of a cell's number format: `FORMAT_NONE` for ordinary numbers, +`FORMAT_DATE` for date-only formats, `FORMAT_TIME` for time-only formats and +`FORMAT_DATETIME` for formats with both date and time components. +""" +@enum CellFormatKind begin + FORMAT_NONE + FORMAT_DATE + FORMAT_TIME + FORMAT_DATETIME +end # Builtin number format ids that denote dates or times (same set xlrd uses): -# 14-22 date/time, 27-36 East Asian date, 45-47 elapsed time, 50-58 East -# Asian date variants. -const BUILTIN_DATE_FORMAT_IDS = Set{Int}([14:22; 27:36; 45:47; 50:58]) +# 14-17 dates, 18-21 times, 22 date+time, 27-36 East Asian dates, 45-47 +# elapsed times, 50-58 East Asian date variants. +const BUILTIN_FORMAT_KINDS = Dict{Int,CellFormatKind}( + Dict(i => FORMAT_DATE for i in [14:17; 27:36; 50:58])..., + Dict(i => FORMAT_TIME for i in [18:21; 45:47])..., + 22 => FORMAT_DATETIME, +) -# Decide whether a custom number format string denotes a date/time. Quoted -# literals, escaped characters, padding/fill markers and bracket sections -# (colors like [Red], conditions like [<=100]) must be ignored; elapsed-time -# tokens like [h] or [ss] count as time. What remains is a date/time format -# iff it contains any of the date/time format codes y, m, d, h or s. -function is_date_format_string(fmt::AbstractString) +# Classify a custom number format string. Quoted literals, escaped characters, +# padding/fill markers and bracket sections (colors like [Red], conditions +# like [<=100]) must be ignored; elapsed-time tokens like [h] or [ss] count as +# time. In what remains, y and d denote a date component, h and s a time +# component, and m either of the two: minutes when next to an h or s code, +# months otherwise. +function format_string_kind(fmt::AbstractString) stripped = IOBuffer() i = firstindex(fmt) n = lastindex(fmt) @@ -36,27 +56,45 @@ function is_date_format_string(fmt::AbstractString) i = nextind(fmt, i) end end - return occursin(r"[ymdhs]"i, String(take!(stripped))) + codes = String(take!(stripped)) + + # AM/PM markers denote a time component but their m is not a month. + hastime = occursin(r"AM/PM|A/P"i, codes) + codes = replace(codes, r"AM/PM|A/P"i => "") + + hastime |= occursin(r"[hs]"i, codes) + hasdate = occursin(r"[yd]"i, codes) + # A run of m codes not adjacent to an h or s code means months, not + # minutes. The lookarounds exclude m itself so that only complete runs + # are considered. + hasdate |= occursin(r"(? "")) + + hasdate && hastime && return FORMAT_DATETIME + hasdate && return FORMAT_DATE + hastime && return FORMAT_TIME + return FORMAT_NONE end -function is_date_format(index::Integer, custom_formats::Dict{UInt16,String}) +function format_kind(index::Integer, custom_formats::Dict{UInt16,String}) # A FORMAT record can redefine any index, including builtin ones, so the # formats stored in the file take precedence over the builtin table. - haskey(custom_formats, index) && return is_date_format_string(custom_formats[index]) - return Int(index) in BUILTIN_DATE_FORMAT_IDS + haskey(custom_formats, index) && return format_string_kind(custom_formats[index]) + return get(BUILTIN_FORMAT_KINDS, Int(index), FORMAT_NONE) end const MILLISECONDS_PER_DAY = 86_400_000 """ - excel_serial_to_temporal(value, is1904) + excel_serial_to_temporal(value, kind, is1904) -Convert an Excel serial date/time value to a `DateTime`, or to a `Time` when -the value has no date component (`0 <= value < 1`). In the 1900 date system -the nonexistent date 1900-02-29 (serial 60) maps to 1900-02-28. Negative -values are not valid dates and are returned unchanged as `Float64`. +Convert an Excel serial date/time value to a `Date`, `DateTime` or `Time`, +depending on the format `kind` of the cell and the value: a date-only format +with no fractional part gives a `Date`, a time-only format with a value below +one day gives a `Time`, and everything else gives a `DateTime`. In the 1900 +date system the nonexistent date 1900-02-29 (serial 60) maps to 1900-02-28. +Negative values are not valid dates and are returned unchanged as `Float64`. """ -function excel_serial_to_temporal(value::Float64, is1904::Bool) +function excel_serial_to_temporal(value::Float64, kind::CellFormatKind, is1904::Bool) value < 0 && return value days = floor(Int, value) ms = round(Int, (value - days) * MILLISECONDS_PER_DAY) @@ -65,7 +103,7 @@ function excel_serial_to_temporal(value::Float64, is1904::Bool) ms = 0 end t = Time(0) + Millisecond(ms) - days == 0 && return t + days == 0 && kind != FORMAT_DATE && return t if is1904 d = Date(1904, 1, 1) + Day(days) elseif days == 60 @@ -75,5 +113,6 @@ function excel_serial_to_temporal(value::Float64, is1904::Bool) else d = Date(1899, 12, 30) + Day(days) end + kind == FORMAT_DATE && ms == 0 && return d return DateTime(d, t) end diff --git a/src/types.jl b/src/types.jl index 454692b..4699204 100644 --- a/src/types.jl +++ b/src/types.jl @@ -50,10 +50,10 @@ mutable struct Workbook <: AbstractWorkbook sheets_info::Vector{WorksheetInfo} sheetname_index::Dict{String,Int} sheets::Dict{Int,Worksheet} - xf_isdate::Vector{Bool} # per XF record: does its number format denote a date/time? + xf_kind::Vector{CellFormatKind} # per XF record: does its number format denote a date and/or time? - function Workbook(handle::Ptr{xlsWorkBook}, is1904::Bool, charset::String, sheets_info::Vector{WorksheetInfo}, sheetname_index::Dict{String,Int}, sheets::Dict{Int,Worksheet}, xf_isdate::Vector{Bool}) - new_wb = new(handle, is1904, charset, sheets_info, sheetname_index, sheets, xf_isdate) + function Workbook(handle::Ptr{xlsWorkBook}, is1904::Bool, charset::String, sheets_info::Vector{WorksheetInfo}, sheetname_index::Dict{String,Int}, sheets::Dict{Int,Worksheet}, xf_kind::Vector{CellFormatKind}) + new_wb = new(handle, is1904, charset, sheets_info, sheetname_index, sheets, xf_kind) finalizer(close, new_wb) return new_wb end diff --git a/src/workbook.jl b/src/workbook.jl index 0cb9fe8..9218ed2 100644 --- a/src/workbook.jl +++ b/src/workbook.jl @@ -28,15 +28,15 @@ function Workbook(filepath::AbstractString) end end - xf_isdate = Vector{Bool}(undef, xlswb.xfs.count) + xf_kind = Vector{CellFormatKind}(undef, xlswb.xfs.count) for i in 1:xlswb.xfs.count xf_data = unsafe_load(xlswb.xfs.xf, i) - xf_isdate[i] = is_date_format(xf_data.format, custom_formats) + xf_kind[i] = format_kind(xf_data.format, custom_formats) end charset = xlswb.charset == C_NULL ? "" : unsafe_string(xlswb.charset) - return Workbook(handle, xlswb.is1904 != 0, charset, sheets_info, sheetname_index, Dict{Int,Worksheet}(), xf_isdate) + return Workbook(handle, xlswb.is1904 != 0, charset, sheets_info, sheetname_index, Dict{Int,Worksheet}(), xf_kind) end """ @@ -79,6 +79,13 @@ function check_xls_file_format(filepath::AbstractString) end end +""" + close(wb::Workbook) + +Close the workbook and release the resources held by the C library. Closing +an already closed workbook does nothing. Worksheets obtained from the +workbook must not be accessed afterwards. +""" function Base.close(wb::Workbook) if wb.handle != C_NULL for ws in values(wb.sheets) @@ -90,10 +97,34 @@ function Base.close(wb::Workbook) return nothing end +""" + isopen(wb::Workbook) + +Whether the workbook has not been closed yet. +""" Base.isopen(wb::Workbook) = wb.handle != C_NULL +""" + sheetcount(wb::Workbook) + +The number of sheets in the workbook, including hidden and empty ones. +""" sheetcount(wb::Workbook)::Int = length(wb.sheets_info) +""" + is1904(wb::Workbook) + +Whether the workbook uses the 1904 date system (the default on classic Mac +versions of Excel) rather than the 1900 date system. Cell values already +account for this, so this is informational only. +""" is1904(wb::Workbook)::Bool = wb.is1904 +""" + sheetname(wb::Workbook, sheet_index) + sheetname(ws::Worksheet) + +The name of the sheet with the given (1-based) index, or of the given +worksheet. +""" sheetname(wb::Workbook, sheet_index::Integer)::String = wb.sheets_info[sheet_index].name @inline is_valid_sheetindex(wb::Workbook, sheet_index::Integer) = 0 < sheet_index <= sheetcount(wb) @@ -105,15 +136,38 @@ end is_valid_sheetname(wb, sheet_name) || error("$sheet_name is not a valid sheet name.") end +""" + sheetindex(wb::Workbook, sheet_name) + sheetindex(ws::Worksheet) + +The (1-based) index of the sheet with the given name, or of the given +worksheet. +""" @inline function sheetindex(wb::Workbook, sheet_name::AbstractString)::Int check_valid_sheetname(wb, sheet_name) return wb.sheetname_index[sheet_name] end +""" + sheetnames(wb::Workbook) + +The names of all sheets in the workbook, in sheet order. +""" sheetnames(wb::Workbook)::Vector{String} = [sheetname(wb, i) for i in 1:sheetcount(wb)] +""" + isvisible(wb::Workbook, sheet_index_or_name) + +Whether the sheet with the given index or name is visible (not hidden). +""" isvisible(wb::Workbook, sheet_index::Integer)::Bool = wb.sheets_info[sheet_index].isvisible isvisible(wb::Workbook, sheet_name::AbstractString)::Bool = isvisible(wb, sheetindex(wb, sheet_name)) +""" + getworksheet(wb::Workbook, sheet_index_or_name) -> Worksheet + +Return the worksheet with the given (1-based) index or name, parsing it on +first access. `wb[i]` and `wb["name"]` are shorthands for this function. +""" function getworksheet(wb::Workbook, sheet_index::Integer)::Worksheet wb.handle == C_NULL && error("Workbook is closed.") if sheet_index ∉ keys(wb.sheets) diff --git a/src/worksheet.jl b/src/worksheet.jl index 780514d..f55eb2a 100644 --- a/src/worksheet.jl +++ b/src/worksheet.jl @@ -20,6 +20,11 @@ function Base.close(ws::Worksheet) return nothing end +""" + size(ws::Worksheet[, dim]) + +The dimensions of the used cell range of the worksheet, as (rows, columns). +""" Base.size(ws::Worksheet) = (ws.lastrow, ws.lastcol) Base.size(ws::Worksheet, dim::Integer) = size(ws)[dim] @@ -82,8 +87,11 @@ cell_marker(cell::st_cell_data) = cell.str == C_NULL ? "" : unsafe_string(cell.s function number_value(wb::Workbook, cell::st_cell_data) xf_index = Int(cell.xf) + 1 - if xf_index <= length(wb.xf_isdate) && wb.xf_isdate[xf_index] - return excel_serial_to_temporal(cell.d, wb.is1904) + if xf_index <= length(wb.xf_kind) + kind = wb.xf_kind[xf_index] + if kind != FORMAT_NONE + return excel_serial_to_temporal(cell.d, kind, wb.is1904) + end end return cell.d end diff --git a/test/test_libxls.jl b/test/test_libxls.jl index 2bcfa13..6ceb338 100644 --- a/test/test_libxls.jl +++ b/test/test_libxls.jl @@ -1,30 +1,114 @@ +@testitem "C struct layout" begin + # The structs in c.jl mirror libxls's structs byte for byte; a mismatch + # means unsafe_load reads garbage. These are the sizes of the structs as + # compiled on all 64-bit platforms (idea from PR #15). 32-bit ABIs differ + # between operating systems, so no fixed sizes are asserted there. + if Sys.WORD_SIZE == 64 + @test sizeof(LibXLS.st_sheet_data) == 16 + @test sizeof(LibXLS.st_sheet) == 16 + @test sizeof(LibXLS.st_font_data) == 24 + @test sizeof(LibXLS.st_font) == 16 + @test sizeof(LibXLS.st_format_data) == 16 + @test sizeof(LibXLS.st_format) == 16 + @test sizeof(LibXLS.st_xf_data) == 24 + @test sizeof(LibXLS.st_xf) == 16 + @test sizeof(LibXLS.str_sst_string) == 8 + @test sizeof(LibXLS.st_sst) == 32 + @test sizeof(LibXLS.st_cell_data) == 40 + @test sizeof(LibXLS.st_cell) == 16 + @test sizeof(LibXLS.st_row_data) == 32 + @test sizeof(LibXLS.st_row) == 16 + @test sizeof(LibXLS.st_colinfo_data) == 10 + @test sizeof(LibXLS.st_colinfo) == 16 + @test sizeof(LibXLS.xlsWorkBook) == 168 + @test sizeof(LibXLS.xlsWorkSheet) == 48 + end + + # The version of the loaded C library must be the one the struct mirrors + # were written against. + @test startswith(LibXLS.xls_getVersion(), "1.6.") +end + @testitem "Error and format helpers" begin for (code, text) in Dict(0x00 => "#NULL!", 0x07 => "#DIV/0!", 0x17 => "#REF!", 0x2A => "#N/A", 0x1D => "#NAME?", 0x24 => "#NUM!", 0x0F => "#VALUE!") @test sprint(show, CellError(code)) == text end @test sprint(show, CellError(0x99)) == "#ERROR(153)!" - @test LibXLS.is_date_format_string("yyyy-mm-dd") - @test LibXLS.is_date_format_string("h:mm AM/PM") - @test LibXLS.is_date_format_string("[h]:mm:ss") - @test LibXLS.is_date_format_string("[Red]dd/mm/yyyy") - @test !LibXLS.is_date_format_string("General") - @test !LibXLS.is_date_format_string("0.00") - @test !LibXLS.is_date_format_string("#,##0.00") - @test !LibXLS.is_date_format_string("0.00E+00") - @test !LibXLS.is_date_format_string("\"m\"0.00") - @test !LibXLS.is_date_format_string("[Red]0.00") + using LibXLS: format_string_kind, format_kind, FORMAT_NONE, FORMAT_DATE, FORMAT_TIME, FORMAT_DATETIME + + @test format_string_kind("yyyy-mm-dd") == FORMAT_DATE + @test format_string_kind("dd/mm") == FORMAT_DATE + @test format_string_kind("mmm") == FORMAT_DATE + @test format_string_kind("[Red]dd/mm/yyyy") == FORMAT_DATE + @test format_string_kind("h:mm") == FORMAT_TIME + @test format_string_kind("h:mm AM/PM") == FORMAT_TIME + @test format_string_kind("hh:mm:ss.000") == FORMAT_TIME + @test format_string_kind("[h]:mm:ss") == FORMAT_TIME + @test format_string_kind("mm:ss") == FORMAT_TIME + @test format_string_kind("yyyy-mm-dd hh:mm") == FORMAT_DATETIME + @test format_string_kind("d-mmm h:mm AM/PM") == FORMAT_DATETIME + @test format_string_kind("General") == FORMAT_NONE + @test format_string_kind("0.00") == FORMAT_NONE + @test format_string_kind("#,##0.00") == FORMAT_NONE + @test format_string_kind("0.00E+00") == FORMAT_NONE + @test format_string_kind("\"m\"0.00") == FORMAT_NONE + @test format_string_kind("\"unterminated") == FORMAT_NONE + @test format_string_kind("[Red]0.00") == FORMAT_NONE + @test format_string_kind("[unterminated") == FORMAT_NONE + @test format_string_kind("\\y0.00") == FORMAT_NONE + @test format_string_kind("_y0.00") == FORMAT_NONE + @test format_string_kind("*y0.00") == FORMAT_NONE + + # builtin format ids, and FORMAT records taking precedence over them + @test format_kind(14, Dict{UInt16,String}()) == FORMAT_DATE + @test format_kind(19, Dict{UInt16,String}()) == FORMAT_TIME + @test format_kind(22, Dict{UInt16,String}()) == FORMAT_DATETIME + @test format_kind(46, Dict{UInt16,String}()) == FORMAT_TIME + @test format_kind(0, Dict{UInt16,String}()) == FORMAT_NONE + @test format_kind(164, Dict{UInt16,String}(0x00a4 => "yyyy-mm-dd")) == FORMAT_DATE + @test format_kind(14, Dict{UInt16,String}(0x000e => "0.00")) == FORMAT_NONE using Dates - @test LibXLS.excel_serial_to_temporal(42066.0, false) == DateTime(2015, 3, 3) - @test LibXLS.excel_serial_to_temporal(42039.4263888889, false) == DateTime(2015, 2, 4, 10, 14) - @test LibXLS.excel_serial_to_temporal(32242.0, false) == DateTime(1988, 4, 9) - @test LibXLS.excel_serial_to_temporal(0.626388888888889, false) == Time(15, 2, 0) - @test LibXLS.excel_serial_to_temporal(1.0, false) == DateTime(1900, 1, 1) - @test LibXLS.excel_serial_to_temporal(59.0, false) == DateTime(1900, 2, 28) - @test LibXLS.excel_serial_to_temporal(60.0, false) == DateTime(1900, 2, 28) # Excel's phantom 1900-02-29 - @test LibXLS.excel_serial_to_temporal(61.0, false) == DateTime(1900, 3, 1) - @test LibXLS.excel_serial_to_temporal(1.0, true) == DateTime(1904, 1, 2) + using LibXLS: excel_serial_to_temporal + + @test excel_serial_to_temporal(42066.0, FORMAT_DATE, false) == Date(2015, 3, 3) + @test excel_serial_to_temporal(42066.0, FORMAT_DATE, false) isa Date + @test excel_serial_to_temporal(42066.0, FORMAT_DATETIME, false) == DateTime(2015, 3, 3) + @test excel_serial_to_temporal(42066.0, FORMAT_DATETIME, false) isa DateTime + @test excel_serial_to_temporal(42039.4263888889, FORMAT_DATETIME, false) == DateTime(2015, 2, 4, 10, 14) + @test excel_serial_to_temporal(42039.4263888889, FORMAT_DATE, false) == DateTime(2015, 2, 4, 10, 14) + @test excel_serial_to_temporal(0.626388888888889, FORMAT_TIME, false) == Time(15, 2, 0) + @test excel_serial_to_temporal(0.626388888888889, FORMAT_DATETIME, false) == Time(15, 2, 0) + @test excel_serial_to_temporal(1.5416666666, FORMAT_TIME, false) == DateTime(1900, 1, 1, 13, 0) + @test excel_serial_to_temporal(1.0, FORMAT_DATE, false) == Date(1900, 1, 1) + @test excel_serial_to_temporal(59.0, FORMAT_DATE, false) == Date(1900, 2, 28) + @test excel_serial_to_temporal(60.0, FORMAT_DATE, false) == Date(1900, 2, 28) # Excel's phantom 1900-02-29 + @test excel_serial_to_temporal(61.0, FORMAT_DATE, false) == Date(1900, 3, 1) + @test excel_serial_to_temporal(1.0, FORMAT_DATE, true) == Date(1904, 1, 2) + @test excel_serial_to_temporal(-1.0, FORMAT_DATE, false) == -1.0 # not a valid date + # sub-millisecond rounding carries over into the next day + @test excel_serial_to_temporal(0.99999999999, FORMAT_TIME, false) == DateTime(1900, 1, 1) + + # C error reporting + @test LibXLS.xls_getError(LibXLS.LIBXLS_OK) isa String + @test !isempty(LibXLS.xls_getError(LibXLS.LIBXLS_ERROR_PARSE)) + @test LibXLS.expect(LibXLS.LIBXLS_OK, "fine") === nothing + @test_throws ErrorException LibXLS.expect(LibXLS.LIBXLS_ERROR_READ, "boom") + err = try + LibXLS.expect(LibXLS.LIBXLS_ERROR_READ, "boom") + catch e + e + end + @test occursin("boom", err.msg) +end + +@testitem "Corrupt files" begin + # A file with a valid OLE2 header that libxls cannot parse: the open + # itself must fail with the error reported by the C library. + path = joinpath(mktempdir(), "corrupt.xls") + write(path, vcat(LibXLS.XLS_FILE_HEADER, rand(UInt8, 100))) + @test_throws ErrorException openxls(path) end @testitem "Reading TestData.xls" begin @@ -42,11 +126,15 @@ end @test_throws ErrorException LibXLS.sheetindex(wb, "No Such Sheet") @test !LibXLS.is1904(wb) @test LibXLS.isvisible(wb, 1) + @test LibXLS.isvisible(wb, "Sheet1") + @test sprint(show, wb) == "LibXLS.Workbook with 4 sheet(s)" ws = getworksheet(wb, "Sheet1") @test ws === wb["Sheet1"] === wb[1] @test size(ws, 1) >= 7 @test size(ws, 2) >= 14 + @test size(ws) == (size(ws, 1), size(ws, 2)) + @test sprint(show, ws) == "LibXLS.Worksheet Sheet1 ($(size(ws, 1))x$(size(ws, 2)))" # Row 3 holds the headers, data starts at row 4; columns C (3) onwards. @test ws[1, 1] === missing @@ -59,10 +147,13 @@ end @test ws[4, 5] === true @test ws[5, 5] isa Bool @test ws[6, 7] === missing # "Mixed with NA" NA cell - @test ws[4, 11] == DateTime(2015, 3, 3) # "Some dates" + @test ws[4, 11] == Date(2015, 3, 3) # "Some dates" + @test ws[4, 11] isa Date @test ws[5, 11] == DateTime(2015, 2, 4, 10, 14) - @test ws[6, 11] == DateTime(1988, 4, 9) + @test ws[5, 11] isa DateTime + @test ws[6, 11] == Date(1988, 4, 9) @test ws[7, 11] == Time(15, 2, 0) + @test ws[7, 11] isa Time @test ws[5, 12] == DateTime(1950, 8, 9, 18, 40) @test ws[7, 12] === missing # "Dates with NA" NA cell @test ws[4, 13] isa CellError # "Some errors" @@ -81,9 +172,12 @@ end @test ws2[10, 6] === false @test ws2[13, 9] == Time(15, 2, 0) + @test isopen(wb) close(wb) + close(wb) # closing twice is fine @test !isopen(wb) @test_throws ErrorException getworksheet(wb, 1) + @test_throws ErrorException ws[4, 3] # do-block form closes the workbook result = openxls(filename) do wb2 @@ -138,12 +232,13 @@ end ws = wb["Plan1"] @test ws[2, 2] == 1 + @test ws[2, 5] isa Date check_test_data(ws, [ [missing for i in 1:6], [missing, 1, 2, 3, missing, 5], [missing, 1000.1, 1000.2, 1000.3, missing, 1000.5], [missing, "abc", "def", "ghi", missing, "xyz"], - [missing, DateTime(2018, 12, 1), DateTime(2018, 12, 31), DateTime(2019, 1, 1), missing, DateTime(2019, 2, 26)], + [missing, Date(2018, 12, 1), Date(2018, 12, 31), Date(2019, 1, 1), missing, Date(2019, 2, 26)], ]) ws2 = wb["Plan2"] @@ -163,7 +258,7 @@ end @test LibXLS.is1904(wb) @test sheetnames(wb) == ["Plan1", "Plan2"] ws = wb["Plan1"] - @test ws[2, 5] == DateTime(2018, 12, 1) - @test ws[4, 5] == DateTime(2019, 1, 1) + @test ws[2, 5] == Date(2018, 12, 1) + @test ws[4, 5] == Date(2019, 1, 1) end end