Mistral 4: a quick RE benchmark of "lechonk"

Some context

Mistral 4 large model (aka "le chonk") was released the 6th of October. It's a great news since it's the best European model to date and it's described as "one of the world's strongest AI models for cybersecurity". So of course it caught our interest, in particular the post below:

Mistral 4 large can handle a plain-text cobalt strike beacon in 12 minutes, wahou
Figure 2: Mistral 4 large can handle a plain-text cobalt strike beacon in 12 minutes, wahou

Now if you ask me, they missed their PR on this one. Detecting a plain text Cobalt Strike in 12 minutes is far from impressive, a human would be faster. And it does not tell a lot: all malware are obfuscated, who cares about a plain-text beacon?

Alas, since I've already benchmarked several models in April, I thought I should stop complaining and do a quick benchmark of "le chonk". Since April, I have improved Malcat's automated malware analysis pipeline (named Malcat Logos, still in beta) a lot AND models have improved as well. So it would be a good occasion to revisit these benchmarks. I will be honest, as a french dev I kind of cheer for Mistral and hope it will be up to the task. Let's see!

TL;DR: If you are in a hurry, you can jump to the result of the benchmark here.

The benchmark

My interest is in Mistral 4's reverse engineering capabilities. Assessing reverse engineering capabilities across models is not a simple task. Imho a good benchmark to test this capability is the static unpacking benchmark: give a LLM access to Malcat's MCP (and only Malcat MCP), and ask it to recover the next stages of a couple of obfuscated malware. This is easy to evaluate (just count how many layers it peeled) and tests various aspects of the model:

  • Its code analysis capabilities: it has to understand how the next stage is packaged/encrypted/compressed by analysing code
  • Its tools call performances: the decryption part, where LLMs have to chain several decryption algorithm, is particularly tricky
  • Its reasoning, since it often has to explore several tracks and try different approaches for the most complex malware

Each sample is scored out of 5, with partial credit for recovered stages and useful findings; any penalties are explained alongside the results.

Since April, I had time to review even more models and I have found a couple that work pretty well for reverse engineering at a reasonable price: mimo 2.6 pro and glm 5.3 (and their flash versions). Since their size is comparable to Mistral 4 large, I thought they would make good contestants for this benchmark. Regarding the test modalities, models will only have the following limits:

  • Maximum 30 minutes. After 30 minutes, Logos's harness will ask for the final report and ignore any following tool call
  • Budget of 1€ per sample (that's why I'm not even testing Opus and Sol :)
  • MCP: Malcat MCP only (including kesakode requests), no bash nor python

Regarding the budget, please note that Mistral 4 Large was 50% off during this test.

These limits are important: nowadays, almost all models are pretty capable, and given an unlimited budget and time they would all come to a solution eventually. The question is: can they do it efficiently? Let us see through 6 examples.

Sample 1

Sample:
674f19126e6dcf0ebb2bf9944841c4cd43195f73b006d51826c7c4252a7e2122 (VT, Bazaar)
Type:
.NET dropper
Description:
.NET string -> base64 -> 3des -> 404 Keylogger

For the first unpacking challenge, I've chosen a relatively simple sample. A .NET dropper embeds a large base64 string which gets 3DES decrypted. The only subtlety there is that the 3DES key and IV are not generated in the same control flow path as the one where the decryption takes place, since the decryption happens in an asynchronous task. So the model will have to correlate both, which is honestly rather easy.

Key and IV are stored in global fields before decryption task is started
Figure 3: Key and IV are stored in global fields before decryption task is started

I expect all models to at least locate the payload, thanks to Malcat's HugeStringBase64 anomaly! Let us see the results

Sample #1 results
Figure 4: Sample #1 results

This first challenge is a walk in the park for the models. What we can already assess is that:

  • Mistral large 4's inference speed seems to be on point: 55 steps in 6 minutes is impressive
  • Its cost is less impressive, by far the most expensive

Sample 2

Sample:
a64e4bbea5983eefb772b8b467504f6242c913e98a8c7fa9a6fd6e4b9a3631de (VT, Bazaar)
Type:
Golang PE
Description:
A simple single-stage rdata dropper, encryption: add 0xc4 + reverse -> Remus

The second unpacking challenge is also rather simple: a Golang malware that drops an encrypted Remus instance from its rdata section. I say easy mostly because:

  • Malcat points directly to the payload with its BigBufferNoXrefMediumToHighEntropy anomaly
  • Not a lot of unknown code to analyse thanks to Kesakode
  • The decryption algorithm is really simple
Easy-peasy right?
Figure 5: Easy-peasy right?

Anyway, let us have a look at the results. You can find Mistral's report here

Sample #2 results
Figure 6: Sample #2 results

Again, pretty easy for all models: they all got it right. But if I were to split hairs:

  • Mistral large 4 was again the most expensive model (don't forget it's with a 50% discount!)
  • Mistral large 4 needed the most steps
  • Just look at mimo 2.6 flash performances!

Sample 3

Sample:
fa755134d9c9796b2f58fd61aeb0ef12121da6afaa1943f05334d332992cdff5 (VT, Bazaar)
Type:
Golang PE
Description:
Golang rdata -> AES -> Donut SC -> ValleyRAT

We stay in Golang territory for this third challenge, but increase the difficulty slightly. The simple XOR is now replaced by AES decryption in ECB mode, which leads to a second stage: a Donut loader. Unpacking the Donut loader gives us the final malware: ValleyRAT. Note that Malcat natively embeds a Donut unpacker, so for this part the model just has to call the right tool.

I really wonder where the payload could be, and where the key is !?!
Figure 7: I really wonder where the payload could be, and where the key is !?!

So let's see if our models perform as well this time!

Sample #3 results
Figure 8: Sample #3 results

This malware was analysed pretty well overall, 5 points for every model. Again Mistral large 4 is the most expensive model, and is saved by its fast inference speed.

Sample 4

Sample:
4109d17d439e425d24e9d11956adcc63ff8e24ccfffe21dd8c5431fe969d2783 (VT, Bazaar)
Type:
PowerShell script
Description:
PowerShell -> Base64 -> Gzip -> PowerShell -> Base64 -> Xor35 -> Cobalt Strike

I don't know about you, but I'm tired of binaries, so let us analyse a small script for a change. Here we have a 3-stage loader:

  • The first stage is a PowerShell script that will base64-decode and gunzip a large string
  • The second stage is another PowerShell script that will base64-decode and XOR a large string
  • The third stage is a Cobalt Strike beacon
Malcat transforms can still be used on text files
Figure 9: Malcat transforms can still be used on text files

Rather easy BUT there is a catch. Malcat is not really well suited to analysing text files. The models will be able to read the content of the text file and call Malcat's transforms, but that's it: no anomalies, no analysis, no Kesakode. They are basically on their own! So let us see how they perform under these conditions:

Sample #4 results
Figure 10: Sample #4 results

In 10 minutes, a mimo 2.6 fast went through 2 layers AND analysed the cobalt strike beacon, I told you that analysing a plain-text beacon in 12 minutes was not impressive :).

The other three models were kind of disappointing for this sample: they got to the Cobalt Strike beacon near the end of their timeout. For Mistral Large 4's defense, it did extract the cobalt strike configuration (see the report) even it it was not part of the prompt. Still 5 points for everyone.

Sample 5

Sample:
ee0f0f2f089ee0594da5750bb4e342c34d703ea045ed80c3b73c81d2f3de8bd4 (VT, Bazaar)
Type:
MSI installer
Description:
MSI -> CAB -> NSIS -> License.txt (obfuscated pyc) -> hex decode -> XOR -> reverse -> b64 -> zlib -> PowerShell -> AES -> some .NET shit I don't remember

Here we face a very long chain: an MSI installer that contains an NSIS installer (and a small start.exe .NET launcher). The NSIS installer packs a decoy installer, a full Python distribution (with a renamed python.exe) and a couple of obfuscated .PYC files named License.txt and License1.txt. The Python bytecode embeds a huge hex-encoded string that gets decrypted to a PowerShell script, which injects a .NET malware that, if I'm being honest, I forgot everything about.

and that's just a small part of the chain :D
Figure 11: and that's just a small part of the chain :D

This will should be tough for our models, even if they can count on Malcat's new python decompiler. But enough talk, let's see the results.

Sample #5 results
Figure 12: Sample #5 results

This one definitely gave more work to the LLMs! Only Mimo 2.6 flash managed to unpack fully the sample up to the c# source (in an impressive 10 minutes run!). For the others, we have partial results. Details are given below:

  • Mistral Large 4: 4/5 (see report):

    • detected the renamed python
    • got through the obfuscated License.txt PYC
    • detected pythonet usage
    • got the CnC url from the PS script
    • did NOT recover the C# source from the powershell script (but got the key and IV right)
  • Mimo v2.6 pro: 3/5 (see report):

    • detected the renamed python
    • got through the obfuscated License.txt PYC
    • got the CnC url
    • did NOT recover the C# source from the powershell script (confused the IV for some "magic prefix")
  • GLM 5.3: 3/5 (see report):

    • detected the renamed python
    • got through the obfuscated License.txt PYC
    • got the CnC url
    • did NOT recover the C# source from the powerhsell script
  • Mimo v2.6 flash: 5/5 (see report):

    • detected the renamed python
    • got through the obfuscated License.txt PYC
    • even deobfuscated the second License1.txt file
    • got the CnC url
    • did recover the C# source from the powershell script

Sample 6

Sample:
33df5d6b61564f89fb40a5c32fc059b1d15fd7971bf3c6db9671100c3ff673f8 (VT, Bazaar)
Type:
Golang loader
Description:
Golang rdata -> custom odd/even arithmetic -> reverse -> byte swap -> xor -> CipherPaneRAT

Here we have a two-stages Golang malware whose difficulty lies in its custom encryption, that needs to be reimplemented using Malcat's sandboxed python transform (a transform that accept a very limited subset of python).

not sure if decompilation really helps :D
Figure 13: not sure if decompilation really helps :D

Another difficulty: the decryption happens in main.main, a very obfuscated 82Kb method! So here we will also judge how models can handle very large amounts of code. But let's see the results:

Sample #6 results
Figure 14: Sample #6 results

This time, theit was the budget limit that stopped Mistral large 4 in its track. According to its report, "le chonk" did simply gave up when faced with the obfuscated large functions, and will be awarded 0 point.

Note: Malcat's MCP allows LLMs to disassemble and decompile page by page

Of the 4 models, only mimo 2.6 flash managed to pierce through the obfuscation and will get all 5 points (see its report). GLM 5.3 flash did recover most of the transforms up to the byte swap, and will be awarded 2 points (I deduced 1 point because the final rapport generation took almost 20 minutes!). Finally mimo 2.6 pro also failed completely and will get 0 point too.

Conclusion

This small benchmark reached its end. It is not a broad benchmark of Mistral 4 large, just a quick test to see how it fits my particular workflow. Regarding raw performances, I think it performed relatively well, on the level of a Mimo v2.6 pro or a GLM 5.3 flash. Its very fast inference speed allowed it to reach this performance even faster.

Its main downside is its price though: at full price, these runs would cost about 12 to 27 times as much as its competitors, without a significant gain in score. Here are the totals across the six samples:

Model Total points Total time Total cost Inference speed (steps/minute) Points / €
Mimo 2.6 flash 30/30 49m 17s ≈ €0.37 3.79 ≈ 81.08
GLM 5.3 flash 25/30 2h 29m 34s ≈ €0.64 1.77 ≈ 39.06
Mistral 4 large 24/30 1h 58m 18s ≈ €9.84 2.39 ≈ 2.44
Mimo 2.6 pro 23/30 2h 17m 50s ≈ €0.80 1.35 ≈ 28.75

Mistral's cost is extrapolated at full price by doubling the discounted amount shown in the screenshots; scores and durations are from the original runs.

The real winner is my old favorite: mimo 2.6 flash, the only model to get everything right, which was also by far the fastest one and also the cheapest one. An explanation could be that speed is an important factor in this kind of time-limited benchmark. But to be honest, "le chonk" was fast, just not good enough.

At the end I would need to throw more complicated samples at these 4 models to really highlights their differences. But I fear Mistral 4 Large's high price point will always weight it down at the end.