Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
301 changes: 150 additions & 151 deletions Doc/c-api/unicode.rst
Original file line number Diff line number Diff line change
Expand Up @@ -154,29 +154,6 @@ access to internal read-only data of Unicode objects:
.. versionadded:: 3.3


.. c:function:: void PyUnicode_WRITE(int kind, void *data, \
Py_ssize_t index, Py_UCS4 value)

Write the code point *value* to the given zero-based *index* in a string.

The *kind* value and *data* pointer must have been obtained from a
string using :c:func:`PyUnicode_KIND` and :c:func:`PyUnicode_DATA`
respectively. You must hold a reference to that string while calling
:c:func:`!PyUnicode_WRITE`. All requirements of
:c:func:`PyUnicode_WriteChar` also apply.

The function performs no checks for any of its requirements,
and is intended for usage in loops.

While :class:`str` objects are usually immutable in Python, this special C API allows
mutating a fresh :class:`str` object if the string has not been "used" yet.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: Py_UCS4 PyUnicode_READ(int kind, void *data, Py_ssize_t index)

Read a code point from a canonical representation *data* (as obtained with
Expand Down Expand Up @@ -385,44 +362,6 @@ Creating and accessing Unicode strings
To create Unicode objects and access their basic sequence properties, use these
APIs:

.. c:function:: PyObject* PyUnicode_New(Py_ssize_t size, Py_UCS4 maxchar)

Create a new Unicode object. *maxchar* should be the true maximum code point
to be placed in the string. As an approximation, it can be rounded up to the
nearest value in the sequence 127, 255, 65535, 1114111.

On error, set an exception and return ``NULL``.

After creation, the string can be filled by :c:func:`PyUnicode_WriteChar`,
:c:func:`PyUnicode_CopyCharacters`, :c:func:`PyUnicode_Fill`,
:c:func:`PyUnicode_WRITE` or similar.
Since strings are supposed to be immutable, take care to not “use” the
result while it is being modified. In particular, before it's filled
with its final contents, a string:

- must not be hashed,
- must not be :c:func:`converted to UTF-8 <PyUnicode_AsUTF8AndSize>`,
or another non-"canonical" representation,
- must not have its reference count changed,
- must not be shared with code that might do one of the above.

This list is not exhaustive. Avoiding these uses is your responsibility;
Python does not always check these requirements.

To avoid accidentally exposing a partially-written string object, prefer
using the :c:type:`PyUnicodeWriter` API, or one of the ``PyUnicode_From*``
functions below.

While :class:`str` objects are usually immutable in Python, this special C API
returns a :class:`str` object that can be mutated, except if *size* is zero, in which
case it returns the immutable empty string constant.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: PyObject* PyUnicode_FromKindAndData(int kind, const void *buffer, \
Py_ssize_t size)

Expand Down Expand Up @@ -755,96 +694,6 @@ APIs:
.. versionadded:: 3.3


.. c:function:: Py_ssize_t PyUnicode_CopyCharacters(PyObject *to, \
Py_ssize_t to_start, \
PyObject *from, \
Py_ssize_t from_start, \
Py_ssize_t how_many)

Copy characters from one Unicode object into another. This function performs
character conversion when necessary and falls back to :c:func:`!memcpy` if
possible. Returns ``-1`` and sets an exception on error, otherwise returns
the number of copied characters.

While :class:`str` objects are usually immutable in Python, this special C API allows
mutating a fresh :class:`str` object if the string has not been "used" yet.

See :c:func:`PyUnicode_New` for details.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: int PyUnicode_Resize(PyObject **unicode, Py_ssize_t length);

Resize a Unicode object *\*unicode* to the new *length* in code points.

Try to resize the string in place (which is usually faster than allocating
a new string and copying characters), or create a new string.

*\*unicode* is modified to point to the new (resized) object and ``0`` is
returned on success. Otherwise, ``-1`` is returned and an exception is set,
and *\*unicode* is left untouched.

The function doesn't check string content, the result may not be a
string in canonical representation.

While :class:`str` objects are usually immutable in Python, this special C API
can resize a :class:`str` object in-place if the string has not been "used" yet.
It returns a :class:`str` object which can be mutated, except if *size* is zero, in
which case it returns the immutable empty string constant.

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: Py_ssize_t PyUnicode_Fill(PyObject *unicode, Py_ssize_t start, \
Py_ssize_t length, Py_UCS4 fill_char)

Fill a string with a character: write *fill_char* into
``unicode[start:start+length]``.

Fail if *fill_char* is bigger than the string maximum character, or if the
string has more than 1 reference.

Return the number of written characters, or return ``-1`` and raise an
exception on error.

While :class:`str` objects are usually immutable in Python, this special C API allows
mutating a fresh :class:`str` object if the string has not been "used" yet.

See :c:func:`PyUnicode_New` for details.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: int PyUnicode_WriteChar(PyObject *unicode, Py_ssize_t index, \
Py_UCS4 character)

Write a *character* to the string *unicode* at the zero-based *index*.
Return ``0`` on success, ``-1`` on error with an exception set.

This function checks that *unicode* is a Unicode object, that the index is
not out of bounds, and that the object's reference count is one.
See :c:func:`PyUnicode_WRITE` for a version that skips these checks,
making them your responsibility.

While :class:`str` objects are usually immutable in Python, this special C API allows
mutating a fresh :class:`str` object if the string has not been "used" yet.

See :c:func:`PyUnicode_New` for details.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: Py_UCS4 PyUnicode_ReadChar(PyObject *unicode, Py_ssize_t index)

Read a character from a string. This function checks that *unicode* is a
Expand Down Expand Up @@ -2053,3 +1902,153 @@ The following API is deprecated.
This API does nothing since Python 3.12.
Previously, this could be called to check if
:c:func:`PyUnicode_READY` is necessary.


.. _pyobject-new-mutating:

Mutating string objects
"""""""""""""""""""""""

The following functions allow creating a blank string object of a given size,
then filling in its contents.
They are :term:`soft deprecated`: they break the assumption that
strings are immutable, making them hard to use correctly.

Prefer using the :c:type:`PyUnicodeWriter` API,
or one of the ``PyUnicode_From*``
functions such as :c:func:`PyUnicode_FromStringAndSize`.

If you do use the functions below, take care to not "use" such a string while
it is being modified.
In particular, before it's filled with its final contents, a string:

- must not be hashed,
- must not be :c:func:`converted to UTF-8 <PyUnicode_AsUTF8AndSize>`,
or another non-"canonical" representation,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Currently, _PyUnicode_IsModifiable() returns 1 even if _PyUnicode_UTF8() is not NULL (for non-ASCII strings). Maybe it would be worth it return 0 in this case.

Note: For compact ASCII strings, _PyUnicode_UTF8() is always non-NULL.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah. If you modify the string after the UTF-8 representation is cached, it'll get out of sync.

For compact ASCII strings, _PyUnicode_UTF8() has undefined behaviour. Maybe it wants an assert.

I filed #159033

- must not have its reference count changed,
Comment thread
encukou marked this conversation as resolved.
- must not be accessed from another thread,
- must not be shared with code that might do one of the above.

This list is not exhaustive. Avoiding these uses is your responsibility;
Python does not always check these requirements.


.. c:function:: PyObject* PyUnicode_New(Py_ssize_t size, Py_UCS4 maxchar)

Create a new Unicode object. *maxchar* should be the true maximum code point
to be placed in the string. As an approximation, it can be rounded up to the
nearest value in the sequence 127, 255, 65535, 1114111.

On error, set an exception and return ``NULL``.

See :ref:`pyobject-new-mutating` for important warnings and caveats.

.. versionadded:: 3.3

.. soft-deprecated:: next
See :ref:`pyobject-new-mutating`.


.. c:function:: void PyUnicode_WRITE(int kind, void *data, \
Py_ssize_t index, Py_UCS4 value)

Write the code point *value* to the given zero-based *index* in a string.

The *kind* value and *data* pointer must have been obtained from a
string using :c:func:`PyUnicode_KIND` and :c:func:`PyUnicode_DATA`
respectively. You must hold a reference to that string while calling
:c:func:`!PyUnicode_WRITE`. All requirements of
:c:func:`PyUnicode_WriteChar` also apply.

The function performs no checks for any of its requirements,
and is intended for usage in loops.

The owning string must not be "used" yet.
See :ref:`pyobject-new-mutating` for details.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: Py_ssize_t PyUnicode_CopyCharacters(PyObject *to, \
Py_ssize_t to_start, \
PyObject *from, \
Py_ssize_t from_start, \
Py_ssize_t how_many)

Copy characters from one Unicode object into another. This function performs
character conversion when necessary and falls back to :c:func:`!memcpy` if
possible. Returns ``-1`` and sets an exception on error, otherwise returns
the number of copied characters.

The destination string must not be "used" yet.
See :ref:`pyobject-new-mutating` for details.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: int PyUnicode_Resize(PyObject **unicode, Py_ssize_t length);

Resize a Unicode object *\*unicode* to the new *length* in code points.

Try to resize the string in place (which is usually faster than allocating
a new string and copying characters), or create a new string.

*\*unicode* is modified to point to the new (resized) object and ``0`` is
returned on success. Otherwise, ``-1`` is returned and an exception is set,
and *\*unicode* is left untouched.

The function doesn't check string content, the result may not be a
string in canonical representation.

*\*unicode* must not be "used" yet.
See :ref:`pyobject-new-mutating` for details.

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: Py_ssize_t PyUnicode_Fill(PyObject *unicode, Py_ssize_t start, \
Py_ssize_t length, Py_UCS4 fill_char)

Fill a string with a character: write *fill_char* into
``unicode[start:start+length]``.

Fail if *fill_char* is bigger than the string maximum character, or if the
string has more than 1 reference.

Return the number of written characters, or return ``-1`` and raise an
exception on error.

*unicode* must not be "used" yet.
See :ref:`pyobject-new-mutating` for details.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: int PyUnicode_WriteChar(PyObject *unicode, Py_ssize_t index, \
Py_UCS4 character)

Write a *character* to the string *unicode* at the zero-based *index*.
Return ``0`` on success, ``-1`` on error with an exception set.

This function checks that *unicode* is a Unicode object, that the index is
not out of bounds, and that the object's reference count is one.
See :c:func:`PyUnicode_WRITE` for a version that skips these checks,
making them your responsibility.

*unicode* must not be "used" yet.
See :ref:`pyobject-new-mutating` for details.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.
Loading