Using Gemma4 with Ollama - Testing File Hash Analysis and Recommendations with AI

Published: 2026-08-12. Last Updated: 2026-08-13 01:26:53 UTC
by Guy Bruneau (Version: 1)
0 comment(s)

In the past few weeks, I have been using Gemma4 as a Large Language Model (LLM) to see how useful it can be to analyze some of the malware hashes uploaded to the DShield sensor over the past 30 days and figure out how its recommendation can be considered useful about the activity my DShield sensor is collecting and tracking. The model I use for this testing is gemma4:e4b [2] using two sites to compare the data against VirusTotal and CyberGordon

The following questions were submitted to Gemma for information:

Instructions:

1. Determine if we should be concerned with the volumes or suspicious volumes match indicators of compromise (IoC) or tactics mentioned on the two external websites.
2. Provide a concise comparative analysis and recommendations for action to prevent actors/bots to successfully compromise the sensor.
3. Identify what type of malware family (Popular threat label) the top 3 files are associated with.
4. Highlight potential issues with the top 3 inbound hashes (files) downloaded by actors/bots logged in the sensor and which sites have the most information about this activity.

Result from the LLM Analysis

Comparative analysis from Gemma 4 :
 As a cybersecurity analyst, I have analyzed the provided data stream concerning files downloaded by actors/bots to the Cowrie sensors.

Analysis of Downloaded Hashes (Top 10 IoCs)

Findings and Analysis

1. Volume vs. IoCs vs. Tactics
The sheer volume of these hashes (126, 110 events) is a major concern. High volumes do not automatically confirm malicious intent, but when coupled with the fact that they are *downloaded* from an interactive sensor like Cowrie, it indicates successful actor persistence and data exfiltration attempts.

Indicators/Tactics

Indicator of Compromise (IoC): The hashes themselves act as IoCs if they resolve to known malware families or C2 infrastructure components. Since the external sites provided no results, the focus must be on the *behavior*—the repeated transfer of these files.
Tactics (MITRE ATT&CK): This activity strongly suggests **Persistence**, **Command and Control (C2)** communication, and potentially **Lateral Movement/Exfiltration**. The actors are not merely querying the system; they are downloading specific files repeatedly.

2. Comparative Analysis & Recommendations

| Aspect | VirusTotal (VT) | CyberGordon (CG) | Conclusion / Action Recommendation |
| :--- | :--- | :--- | :--- |
| Data Depth | High-fidelity, integrated threat intelligence from multiple sources (AV engines). Best for direct identification. | Academic/Historical analysis engine. Useful for deep source tracking and niche observations. | VT is superior for immediate threat context; CG offers supporting research. |
| Reliability | Generally reliable, though sometimes slow or requires paid access for full fidelity. | Utility appears high but the provided output is generic (focusing on JS compatibility rather than results). | **Rely primarily on VT/Commercial Threat Feeds.** Do not treat lack of a result as "safe." |
| Missing Data | The provided snippet was boilerplate ("Please enable JavaScript"), meaning we missed vital IoC intelligence. | The provided snippet was generic browser/engine advice, offering no actionable threat intel. | **Critical Action:** Develop a structured workflow to *ensure* the results page is fully processed and captured for all 10 hashes. |

Recommendations for Action (Preventing Compromise):

1.  Immediate Containment: Treat any machine connected through this sensor as compromised. Isolate the host and initiate forensic imaging immediately.
2.  Mitigation at Sensor Level: Update Cowrie's logging and alerting capabilities to track file *execution* attempts, not just downloads. Implement stricter egress filtering on the network hosting the sensors.
3.  Threat Hunting (Proactive): Utilize the Top 10 hashes to query internal Endpoint Detection and Response (EDR) solutions across the entire enterprise, searching for the hash signatures or file names appearing on any system *outside* of the sensor environment.
4.  Hardening: Review user access policies and network segmentation rules. The repeated downloading suggests a persistent gap in perimeter controls.

3. Associated Malware Family (Top 3 Hashes)

Since no specific threat intelligence results were provided for these hashes, this analysis is based on the pattern of high-volume command-and-control behavior.

1.  197c74408e15bd1168105f564f96aace4fd4819961b724630bf5a6be4878daf8: **Botnet/Loader Malware.** The extremely high event count and consistent downloading pattern are classic indicators of a payload being repeatedly delivered, likely by an established botnet or initial access Trojan (RAT).
2.  31d4181843b1ed10a7e7cb3f108f6d6c50a7a4452ee52ddacabe8ca77260615e: **Backdoor/Keylogger.** Similar to the top hash, but potentially representing a specific piece of payload used for data exfiltration (like credentials or session tokens).
3.  94f2e4d8d4436874785cd14e6e6d403507b8750852f7f2040352069a75da4c00: **Credential Stealer/Dropper.** This hash may represent a modular component or a staged payload, designed to "drop" the final malware onto the compromised system after initial access.

4. Potential Issues with Top 3 Hashes & Best Data Source

Potential Issue Highlight (Focusing on Behavioral Risk)

The primary issue is not necessarily what the hashes *are*, but that they are actively being *downloaded* repeatedly by a controlled actor over a sustained period (Top 10). This behavior confirms **successful compromise and established persistence**. The actors have successfully identified data pathways to download their chosen payloads, bypassing existing security controls.

 

Comparison of Sites for Top 3

Given the critical nature of these files:

VirusTotal (VT): Is the preferred source. Despite the provided boilerplate text, VT's strength lies in its aggregation of multiple vendor signatures (AV engines). It provides the highest likelihood of a quick determination whether the hash is *known malicious*.
CyberGordon (CG): Is useful as a secondary research source or for understanding the historical context of the observed payload type. Its data output, based on observation and academic logging, might reveal more about *how* the file was used in previous compromises, which is valuable for threat hunting.

Conclusion

The most reliable actionable intelligence will be obtained by re-querying both sites multiple times to ensure 100% capture of the IoC data. However, **VT provides the immediate signature risk assessment.**

My own conclusion: I think the information provided by Gemma is a nice summary as long as the right tools (sites) are provided that can be used to gather additional information to use to better answer the questions initially asked in the instructions

The python script used for this test is available here

[1] https://en.wikipedia.org/wiki/Large_language_model
[2] https://ollama.com/library/gemma4
[3] https://cybergordon.com
[4] https://www.virustotal.com
[5] https://dshield.org/diary/Using+AI+Gemma+3+Locally+with+a+Single+CPU/32556
[6] https://github.com/bruneaug/DShield-SIEM
[7] https://github.com/bruneaug/DShield-SIEM/tree/main/Ollama

-----------
Guy Bruneau IPSS Inc.
My GitHub Page
Twitter: GuyBruneau
gbruneau at isc dot sans dot edu

0 comment(s)

Linux Kernel Process Accounting

Published: 2026-08-12. Last Updated: 2026-08-12 14:21:10 UTC
by Johannes Ullrich (Version: 1)
0 comment(s)

A couple of days ago, Xavier posted about Atuin to gain more insight into the command history. Atuin does a great job of better organizing what is usually handled by "bash_history" and collecting meaningful additional data. Our reader David commented that this can also be done quite well with Linux's kernel process accounting feature, and I think he is very right. I really like Linux process accounting for a number of reasons, so here is a quick introduction.

Process accounting is a kernel feature. You will not see a specific process responsible for it. Instead, the "accton" command signals the kernel to start logging process data to a specific location (usually /var/log/account/pacct). Once a process terminates, the kernel will log respective details to the binary log file.

1 - Installation

I don't think process accounting is enabled by default on any Linux system. It does add a little additional overhead, but some users may shy away from it because it requires additional disk writes to collect the information. Memory and other CPUs should not be significantly impacted by process accounting. On my not very busy Proxmox system, it uses about 50 MB/day of disk space. So nothing that should be noticeable for most systems.

Installation usually comes down to installing the respective package for your distribution. On Debian based distributions, it is just

apt install acct

This will typically also configure the startup scripts, but it can't hurt to run

systemctl enable --now acct

That is it. Wait a little bit, and you will see the log. 

2 - How to read the logs

Logs are saved in a binary format. The "lastcomm" command can be used to display the log in a readable format. For example:

ip6tables-save   S     root     __         0.00 secs Wed Aug 12 06:25
iptables-restor  S     root     __         0.00 secs Wed Aug 12 06:25
iptables-save    S     root     __         0.00 secs Wed Aug 12 06:25
check_ssh              100107   __         0.00 secs Wed Aug 12 06:25
cron              F    100000   __         0.00 secs Wed Aug 12 06:25
sh               S     100000   __         0.00 secs Wed Aug 12 06:25
debian-sa1             100000   __         0.00 secs Wed Aug 12 06:25

These are a few lines from my Proxmox server. It logs the process name, Flags (S=super user, F=forked process, D=generated core dump, X=terminated by signal), User name (or ID), CPU execution time, and finally the timestamp at which the process was started. The output may be modified slightly depending on the command-line arguments used.

3 - Remote Logging

Unlike most Linux logs, these logs are not created by syslog. However, you may still read them with syslog to forward them to a central log collector/SIEM. Syslog-ng for example include a "s_pacct" processor for process accounting logs. You enable it with this configuration:

source s_pacct {
    pacct(file("/var/log/account/pacct"));
};

4 - Other useful tools

The "sa" command can be used to easily extract summaries from accounting data. For example, a breakdown by CPU time used by different processes

# sa -c | head -10
  329063  100.00%  259245.37re  100.00%      70.29cp  100.00%         0avio     24320k
     469    0.14%      55.38re    0.02%      52.65cp   74.90%         0avio    124752k   ffmpeg
     488    0.15%       5.43re    0.00%       4.57cp    6.51%         0avio      5338k   apt-get
    2820    0.86%       3.60re    0.00%       3.34cp    4.74%         0avio     13295k   ceph
    1231    0.37%       1.96re    0.00%       1.92cp    2.73%         0avio      1974k   ps
     312    0.09%    1907.83re    0.74%       1.36cp    1.93%         0avio     88176k   named
       7    0.00%   39483.07re   15.23%       0.76cp    1.08%         0avio      6348k   systemd-journal
     123    0.04%   34804.71re   13.43%       0.34cp    0.48%         0avio     20367k   ***other*
       8    0.00%       0.68re    0.00%       0.27cp    0.38%         0avio      3926k   store
     366    0.11%       0.58re    0.00%       0.26cp    0.36%         0avio      2235k   dpkg-deb*

5 - Containers

Process accounting is a kernel feature, and the kernel must be compiled and configured to support process accounting. If you are running Linux containers in Proxmox (the platform I am using), process accounting will not work unless the container is privileged. But it does not have to work. The container processes are logged by the host, which I think is actually better. This way, the logs are more easily centralized, and they can't be tampered with from inside the container.  

6 - Conclusion

I think Linux kernel process accounting is a very neat and often overlooked feature. You may be able to do more fine-grained inspection with eBPF, but process accounting is "ready to go and useful" with little work. It does not log command line options, which may be an issue in incident response. But it is a very good supplement to other features like bash_history files, and it captures processes that bash_history would never see.

 

--
Johannes B. Ullrich, Ph.D. , Dean of Research, SANS.edu
Twitter|

0 comment(s)
ISC Stormcast For Wednesday, August 12th, 2026 https://isc.sans.edu/podcastdetail/10048

Comments


Diary Archives