Repository navigation
TLS/SSL module does not work on its own (Pico w) #9071
Description
Activity
Thanks @Darkness4191 for the detailed report.
My first guess would be something related to allocating buffers in mbedtls, and something about the state the device is in is causing the allocations to fail in some scenarios (e.g. fragmentation). TBH though I would have thought the behaviour would be the opposite way around though.
If this is the case, then it's the rp2040 version of #8940 and the fix is likely the same, although a lot simpler on rp2040 because we don't have to work around the IDF's memory layout.
I will have to try and replicate this locally... here's the likely next steps I would use to investigate:
- Compile the firmware with
MICROPY_GC_HEAP_SIZEfrom 166kiB to (say) 100kiB (see the top of ports/rp2/main.c) to leave more room for the pico-sdk/mbedtls malloc. - Attach SWD debugger and find what line in mbedtls is causing the failure.
ca_data_base64 = "PLACEHOLDER_CA_CERT_DER_BASE64"
ca_data = ubinascii.a2b_base64(ca_data_base64.encode("utf-8"))Just a quick note (and when we write more docs for this we will explain it in more detail) that this way of getting the DER into mbedtls is quite expensive because you end up allocating a lot of string copies. The best thing to do would be to get the DER bytes (in CPython) e.g.
ca_data_base64 = "PLACEHOLDER_CA_CERT_DER_BASE64" ca_data = ubinascii.a2b_base64(ca_data_base64.encode("utf-8")) print(repr(ca_data))
And then use the output
b'...'string directly in your code.(Even better, when we add SSLContext, we're looking at loading the DER files from the filesystem).
- Compile the firmware with
This does not look like an out-of-memory issue, because that should propagate out as a different kind of exception. But you are seeing
MBEDTLS_ERR_X509_CERT_VERIFY_FAILED.It could be the timestamp of the cert. The RTC on the Pico W probably needs to be set to the correct time so the cert validation works. Thonny might be doing this automatically (setting it from the PC's clock).
You can try using this to set the time, once the wifi is connected:
import ntptime ntptime.settime()
Reacted by Mustafa Özgür and PeterThank you @dpgeorge and @jimmo for your quick and detailed answers. I added the time synchronisation dpgeorge suggested and that worked perfectly! This is something I would have never thought of.
I also changed the way the CA certificate is stored by reading it with "rb" from the Pico's flash at runtime.Although this was much fun I would suggest to add a little warning in the documentation or some kind of error because if you need to validate a certificate a correct time is always advantageous.
Again thank you very much for your help. I added these lines of code:
import ntptime ntptime.host = "de.pool.ntp.org" ntptime.settime()Thanks for confirming the problem/fix.
This is definitely a deep trap that is very easy to fall into. Maybe we need to somehow catch this exception and turn it into a nice error message that says something like "certificate is expired, check RTC".
Reacted by beyonlo, ams1 and Hiroki Tsukahara@Darkness4191
Although this was much fun I would suggest to add a little warning in the documentation or some kind of error because if you need to validate a certificate a correct time is always advantageous.
Definitely there should be a note/warning stating that some ports do timestamp cert verification.
Maybe we need to somehow catch this exception and turn it into a nice error message that says something like "certificate is expired, check RTC".
I had a quick look at the mbedtls apidocs and it seems like an easy/ nice idea to implement here in
modussl_mbedtls.ccleanup:cleanup: mbedtls_pk_free(&o->pkey); mbedtls_x509_crt_free(&o->cert); mbedtls_x509_crt_free(&o->cacert); mbedtls_ssl_free(&o->ssl); mbedtls_ssl_config_free(&o->conf); mbedtls_ctr_drbg_free(&o->ctr_drbg); mbedtls_entropy_free(&o->entropy); if (ret == MBEDTLS_ERR_SSL_ALLOC_FAILED) { mp_raise_OSError(MP_ENOMEM); } else if (ret == MBEDTLS_ERR_PK_BAD_INPUT_DATA) { mp_raise_ValueError(MP_ERROR_TEXT("invalid key")); } else if (ret == MBEDTLS_ERR_X509_BAD_INPUT_DATA) { mp_raise_ValueError(MP_ERROR_TEXT("invalid cert")); } else { mbedtls_raise_error(ret); } }
If the error is
MBEDTLS_ERR_X509_CERT_VERIFY_FAILEDthen it could call
mbedtls_x509_crt_verify_info()that should return an informational string about the verification status of a certificate.These are
X509 Verify codes, in this case it would be#define MBEDTLS_X509_BADCERT_EXPIRED 0x01They could be added to
mp_mbedtls_errors.cand maybe only these would be necessary for now.#define MBEDTLS_X509_BADCERT_EXPIRED 0x01 #define MBEDTLS_X509_BADCERT_CN_MISMATCH 0x04 #define MBEDTLS_X509_BADCERT_NOT_TRUSTED 0x08They could be added to mp_mbedtls_errors.c and maybe only these would be necessary for now.
It turns out I wasn't entirely right about this,
to getX509 Verify codes,mbedtls_ssl_get_verify_resultis needed, then the result can be passed tombedtls_x509_crt_verify_info()which gives a nice error string .e.g for a CN mismatch:>>> import test_ssl_context_client Traceback (most recent call last): File "<stdin>", line 1, in <module> File "test_ssl_context_client.py", line 104, in <module> File "test_ssl_context_client.py", line 96, in main File "ssl.py", line 111, in wrap_socket ValueError: The certificate Common Name (CN) does not match with the expected CNAlthough it needs to allocate a buffer for the string, so another option would be manually match the verify codes.
@dpgeorge
I'll push this patch to the draft at #8968, but if you want I could make a PR for currentussl.wrap_socketimplementation 👍🏼Maybe we need to somehow catch this exception and turn it into a nice error message that says something like "certificate is expired, check RTC".
Just a quick thought... The error isn't so much about expiration as it is that the certificate "valid after" date is in the future.
Not trying to parse words, just thought that if you were going to have a developer-friendly error message, it should reflect that the cert isn't valid-after date hasn't been reached yet, not that it has expired.
on it's own
Might you be persuaded to fix this issue's title? The non-grammatical apostrophe is somewhat grating IMO.
Reacted by Robert Quattlebaum- changed the title
[-]TLS/SSL module does not work on it's own (Pico w)[/-][+]TLS/SSL module does not work on its own (Pico w)[/+]on Feb 23, 2023 Commit f3f215e enabled more detailed error messages for SSL certificates. That should help with this problem.
Today I tested the new Raspberry Pico W board and especially the network connectivity. I am aware that every version for that board is marked as "unstable" but I still wanted to Issue a report, because it is a very strange error and I couldn't find any resource on the Internet describing similar issues.
My initial plan was to implement a small IoT service with the Pico W. It should connect to a server with a valid SSL/TLS certificate via TLS1.2, send info about the current power state and also receive commands.
Here is my code for the pico w:
After implementing the backend and writing the pico w script in Thonny everything worked perfectly fine. The CA certificate I hardcoded into the software got accepted and a valid TLS connection could be established. I saved the script as main.py on the pico's flash and triple checked that it actually gets run at startup. I disconnected the pico from my Pc and plugged it into my wall adapter. I waited and I realized the pico wouldn't connect to the server at all. I checked the logs (the server is written in Java) and got "Received fatal alert: certificate_unknown" as the error message on server side. The next thing I did was to connect the pico back to my Pc and watching the execution from another COM reader other than Thonny. It still didn't work as expected with the same error on server side and the error (-9984, 'MBEDTLS_ERR_X509_CERT_VERIFY_FAILED') printed by the Pico.
When I run the script from Thonny I get this output:
When it starts by itself this is the output:
When started over and over by another script this is the output:
First I thought this might happen because Thonny restarts scripts by performing a soft-reset and to reset it normally I would cut the power which is not just a soft reset. Therefore, I implemented a soft reset with sys.exit() when the error happens. But this didn't work either. This documentation is clear about that sys.exit() should perform a soft reset, the same reset I wanted to implement but it just doesn't reset. It just exits to the normal python shell and doesn't restart main.py or boot.py. And even when I do a machine reset and the pico is somehow connected to Thonny it does work. So I assume it is not the machine reset that causes the problem but I am not entirely sure.
After 8 hours of trial and error, I can say most certainly that the Pico W would only establish a valid connection to the server when the program is run or reset by the Thonny IDE. I have no idea why that happens, or how I can solve the problem and I would really appreciate help.
The firmware is rp2-pico-w-20220815-unstable-v1.19.1-284-ga16a330da.uf2