
Some context
The Mistral 4 Large model (aka "le chonk") was released on October, 6th. That's great news, since it's the best European model to date and it's described as "one of the world's strongest AI models for cybersecurity". So of course it caught my interest, particularly the post below:

Now, if you ask me, they missed the mark with their PR on this one. Detecting a plain-text Cobalt Strike beacon in 12 minutes is far from impressive; a human would be faster. And it doesn't tell us much: 99.99% of malware samples are obfuscated, so who cares about a plain-text beacon?
Anyway, since I already benchmarked several models in April, I thought I should stop complaining and do a quick benchmark of "le chonk". Since April, I have made a lot of improvements to Malcat's automated malware analysis pipeline (Malcat Logos, still in beta), AND models have improved as well. So this seemed like a good opportunity to revisit these benchmarks. I'll be honest: as a French dev, I'm kind of rooting for Mistral and hoping it will be up to the task. Let's see!
TL;DR: If you are in a hurry, you can jump to the benchmark results here.
The benchmark
My interest is in Mistral 4's reverse engineering capabilities. Assessing reverse engineering capabilities across models is not a simple task. IMHO, static unpacking is a good way to test these capabilities: give an LLM access to Malcat's MCP (and only Malcat's MCP), and ask it to recover the next stages of a few obfuscated malware samples. This is easy to evaluate (just count how many layers it peels) and tests various aspects of the model:
- Its code analysis capabilities: it has to understand how the next stage is packaged, encrypted or compressed by analysing code.
- Its performance when calling tools: the decryption part, where LLMs have to chain several decryption algorithms, is particularly tricky.
- Its reasoning: it often has to explore different leads and try different approaches for the most complex samples.
Each sample is scored out of 5, with partial credit for recovered stages and useful findings; any penalties are explained alongside the results.
Since April, I've had time to review even more models, and I've found a couple that work pretty well for reverse engineering at a reasonable price: MiMo 2.6 Pro and GLM 5.3 (and their Flash versions) as well as GPT Luna 6. Since they are comparable in size to Mistral 4 Large, I thought they would make good contestants for this benchmark.
For these tests, the models have the following limits:
- Maximum of 30 minutes. After 30 minutes, Logos's harness will ask for the final report and ignore any subsequent tool calls.
- Budget of €1 per sample. That's why I'm not even testing Opus and Sol :)
- MCP: Malcat MCP only (including Kesakode requests), no Bash or Python.
Regarding the budget, please note that Mistral 4 Large was 50% off during this test.
These limits are important: nowadays, almost all models are pretty capable, and given unlimited time and budget, they would all eventually find a solution. The question is: can they do it efficiently? Let's see how they do on six examples.
Sample 1
- Sample:
- 674f19126e6dcf0ebb2bf9944841c4cd43195f73b006d51826c7c4252a7e2122 (VT, Bazaar)
- Type:
- .NET dropper
- Description:
- .NET string -> base64 -> 3DES -> 404 Keylogger
For the first unpacking challenge, I've chosen a relatively simple sample. A .NET dropper embeds a large base64 string that is decrypted using 3DES. The only subtlety is that the 3DES key and IV are not generated in the same control flow path as the decryption code, which runs in an asynchronous task. So the model will have to correlate the two, which is honestly rather easy.

I expect all models to at least locate the payload, thanks to Malcat's HugeStringBase64 anomaly! Let's see the results.

This first challenge is a walk in the park for the models. We can already see that:
- Mistral 4 Large's inference speed seems to be on point: 55 steps in 6 minutes is impressive.
- Its cost is less impressive: it is by far the most expensive model.
Sample 2
- Sample:
- a64e4bbea5983eefb772b8b467504f6242c913e98a8c7fa9a6fd6e4b9a3631de (VT, Bazaar)
- Type:
- Golang PE
- Description:
- A simple single-stage rdata dropper, encryption: add 0xc4 + reverse -> Remus
The second unpacking challenge is also rather simple: a Golang malware sample that drops an encrypted Remus instance from its rdata section. I call it easy mostly because:
- Malcat points directly to the payload with its
BigBufferNoXrefMediumToHighEntropyanomaly. - There isn't much unknown code to analyse, thanks to Kesakode.
- The decryption algorithm is really simple.

Anyway, let's have a look at the results. You can find Mistral's report here.

Again, this was pretty easy for all models: they all got it right. But if I were to split hairs:
- Mistral 4 Large was again the most expensive model (don't forget that's with a 50% discount!).
- Mistral 4 Large needed the most steps.
- Just look at MiMo 2.6 Flash's performance!
Sample 3
- Sample:
- fa755134d9c9796b2f58fd61aeb0ef12121da6afaa1943f05334d332992cdff5 (VT, Bazaar)
- Type:
- Golang PE
- Description:
- Golang rdata -> AES -> Donut SC -> ValleyRAT
We stay in Golang territory for this third challenge, but increase the difficulty slightly. The simple XOR is now replaced by AES decryption in ECB mode, which leads to a second stage: a Donut loader. Unpacking the Donut loader gives us the final malware: ValleyRAT. Note that Malcat has a built-in Donut unpacker, so for this part the model just has to call the right tool.

So let's see if our models perform as well this time!

This malware was analysed pretty well overall: 5 points for every model. Again, Mistral 4 Large is the most expensive model, and is saved by its fast inference speed.
Sample 4
- Sample:
- 4109d17d439e425d24e9d11956adcc63ff8e24ccfffe21dd8c5431fe969d2783 (VT, Bazaar)
- Type:
- PowerShell script
- Description:
- PowerShell -> Base64 -> Gzip -> PowerShell -> Base64 -> Xor35 -> Cobalt Strike
I don't know about you, but I'm tired of binaries, so let us analyse a small script for a change. Here we have a 3-stage loader:
- The first stage is a PowerShell script that will base64-decode and gunzip a large string.
- The second stage is another PowerShell script that will base64-decode and XOR a large string.
- The third stage is a Cobalt Strike beacon.

Rather easy, BUT there's a catch. Malcat is not really well suited to analysing text files. The models will be able to read the contents of the text file and call Malcat's transforms, but that's it: no anomalies, no analysis, no Kesakode. They are basically on their own! So let's see how they perform under these conditions:

In 3 minutes, Luna 6 got through two layers AND analysed the Cobalt Strike beacon. I told you that analysing a plain-text beacon in 12 minutes was not impressive :)
MiMo 2.6 Flash also finished in about 10 minutes.
The other three models were kind of disappointing for this sample: they got to the Cobalt Strike beacon near the end of their time limit. In Mistral 4 Large's defence, it did extract the Cobalt Strike configuration (see the report), even though that wasn't part of the prompt. Still, 5 points for everyone.
Sample 5
- Sample:
- ee0f0f2f089ee0594da5750bb4e342c34d703ea045ed80c3b73c81d2f3de8bd4 (VT, Bazaar)
- Type:
- MSI installer
- Description:
- MSI -> CAB -> NSIS -> License.txt (obfuscated pyc) -> hex decode -> XOR -> reverse -> b64 -> zlib -> PowerShell -> AES -> some .NET shit I don't remember
Here we face a very long chain: an MSI installer that contains an NSIS installer (and a small start.exe .NET launcher). The NSIS installer packs a decoy installer, a full Python distribution (with a renamed python.exe) and a couple of obfuscated .PYC files named License.txt and License1.txt. The Python bytecode embeds a huge hex-encoded string that gets decrypted into a PowerShell script, which injects a .NET malware sample that, if I'm being honest, I've forgotten everything about.

This should be tough for our models, even if they can count on Malcat's new Python decompiler. But enough talk, let's see the results.

This one definitely gave the LLMs more work! Only MiMo 2.6 Flash and Luna 6 managed to fully unpack the sample all the way to the C# source, in about 10 and 12 minutes respectively! For the others, we have partial results. Details are given below:
-
Mistral 4 Large: 4/5 (see report):
- detected the renamed Python executable
- got through the obfuscated
License.txtPYC file - detected pythonnet usage
- got the CnC URL from the PS script
- did NOT recover the C# source from the PowerShell script (but got the key and IV right)
-
MiMo 2.6 Pro: 3/5 (see report):
- detected the renamed Python executable
- got through the obfuscated
License.txtPYC file - got the CnC URL
- did NOT recover the C# source from the PowerShell script (mistook the IV for a "magic prefix")
-
GLM 5.3 Flash: 3/5 (see report):
- detected the renamed Python executable
- got through the obfuscated
License.txtPYC file - got the CnC URL
- did NOT recover the C# source from the PowerShell script
-
MiMo 2.6 Flash: 5/5 (see report):
- detected the renamed Python executable
- got through the obfuscated
License.txtPYC file - even deobfuscated the second file,
License1.txt - got the CnC URL
- did recover the C# source from the PowerShell script
-
Luna 6: 5/5 (see report):
- detected the renamed Python executable
- got through the obfuscated
License.txtPYC file - even deobfuscated the second file,
License1.txt - got the CnC URL
- did recover the C# source from both (!) PowerShell scripts
Sample 6
- Sample:
- 33df5d6b61564f89fb40a5c32fc059b1d15fd7971bf3c6db9671100c3ff673f8 (VT, Bazaar)
- Type:
- Golang loader
- Description:
- Golang rdata -> custom odd/even arithmetic -> reverse -> byte swap -> xor -> CipherPaneRAT
Here we have a two-stage Golang malware sample whose difficulty lies in its custom encryption, which needs to be reimplemented using Malcat's sandboxed Python transform (a transform that accepts a very limited subset of Python).

Another difficulty: the decryption happens in main.main, a very obfuscated 82 KB method! So here we will also judge how well the models handle very large amounts of code. But let's see the results:

This time, it was the budget limit that stopped Mistral 4 Large in its tracks. According to its report, "le chonk" simply gave up when faced with the large, obfuscated functions, and will be awarded 0 points.
Note: Malcat's MCP allows LLMs to disassemble and decompile page by page
Of the five models, only MiMo 2.6 Flash managed to get through the obfuscation and will get all 5 points (see its report). GLM 5.3 Flash did recover most of the transforms up to the byte swap, and will be awarded 2 points (I deducted 1 point because generating the final report took almost 20 minutes!). Finally, MiMo 2.6 Pro and Luna 6 also failed completely and will get 0 points too. Luna 6 did not even try :)
Conclusion
That's it for this small benchmark. It is not a broad benchmark of Mistral 4 Large, just a quick test to see how it fits my particular workflow. In terms of raw performance, I think it performed relatively well, on a similar level to MiMo 2.6 Pro, GLM 5.3 Flash and Luna 6. Its very fast inference speed helped it finish sooner than MiMo 2.6 Pro and GLM 5.3 Flash.
Its main downside is its price though: at full price, these runs would cost about 12 to 38 times as much as runs with its competitors, without a significant gain in score.
Here are the totals across the six samples:
| Model | Total points | Total time | Total cost | Inference speed (steps/minute) | Points / € |
|---|---|---|---|---|---|
| MiMo 2.6 Flash | 30/30 | 49m 17s | 0.37€ | 3.79 | 81.08 |
| Luna 6 | 25/30 | 34m 24s | 0.26€ | 3.55 | 96.15 |
| GLM 5.3 Flash | 25/30 | 2h 29m 34s | 0.64€ | 1.77 | 39.06 |
| Mistral 4 Large | 24/30 | 1h 58m 18s | 9.84€ | 2.39 | 2.44 |
| MiMo 2.6 Pro | 23/30 | 2h 17m 50s | 0.80€ | 1.35 | 28.75 |
Mistral's cost is extrapolated at full price by doubling the discounted amount shown in the screenshots; scores and durations are from the original runs.
The real winner is my old favourite: MiMo 2.6 Flash, the only model to get everything right. Luna 6 was the fastest and cheapest overall, but failed hard on the final Go loader, which lets me think it won't be usable for the most advanced malware. It looks like speed is an important factor in this kind of time-limited benchmark, but it doesn't explain everything. To be honest, "le chonk" was fast too, just not good enough.
I would need to throw more complicated samples at these five models to really highlight their differences. But I fear Mistral 4 Large's high price will always weigh it down.