15 seconds at a time, when the owner activates the feature. The scale of its recall is not meaningfully different than someone scribbling down 15 seconds worth of what they hear someone say (or even just remembering it), and does not pose a greater privacy risk than a scratchpad or human memory.
It’s also much less effective at recording people speaking for long durations than the audio recording feature that every smartphone for the last 20 years has supported.
It also does not identify individual speakers.
There are two features, you have listed one of them.