Documentation
¶
Index ¶
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
func CalcRefHash ¶ added in v0.1.39
func CalcRefHash(outValue *string, keyvals ...interface{})
CalcRefHash exposes the shared SHA1 hashing used for deterministic document ids to other packages (the Elasticsearch writer uses it to key the file<->leak reference documents by file_id + leak_id).
Types ¶
type Credential ¶
type Credential struct {
ID uint `json:"id" gorm:"primarykey"`
FileID uint `json:"file_id" gorm:"index:idx_cred"`
Rule string `json:"rule"`
Time time.Time `json:"time"`
UserDomain string `json:"user_domain"`
Username string `json:"username"`
Password string `json:"password"`
CPF string `json:"cpf"`
Url string `json:"url"`
UrlDomain string `json:"url_domain"`
Severity int `json:"severity"`
Entropy float32 `json:"entropy"`
NearText string `json:"near_text"`
}
func (Credential) LeakDoc ¶ added in v0.1.39
func (cred Credential) LeakDoc() map[string]interface{}
func (Credential) LeakID ¶ added in v0.1.39
func (cred Credential) LeakID() string
func (Credential) LeakType ¶ added in v0.1.39
func (cred Credential) LeakType() string
func (Credential) MarshalJSON ¶
func (cred Credential) MarshalJSON() ([]byte, error)
Custom Marshaller for Credential
func (Credential) RefDoc ¶ added in v0.1.39
func (cred Credential) RefDoc() map[string]interface{}
func (*Credential) Sanitize ¶ added in v0.1.30
func (cred *Credential) Sanitize()
Sanitize removes null bytes and invalid UTF-8 sequences from all string fields
type Document ¶ added in v0.1.36
type Document struct {
ID uint `json:"id" gorm:"primarykey"`
FileID uint `json:"file_id" gorm:"index:idx_document"`
Time time.Time `json:"time"`
Raw string `json:"raw"`
Number string `json:"number"`
IsCPF bool `json:"is_cpf"`
IsCNPJ bool `json:"is_cnpj"`
Source string `json:"source"`
FileName string `json:"file_name"`
Line string `json:"line"`
NearText string `json:"near_text"`
}
func (Document) MarshalJSON ¶ added in v0.1.36
Custom Marshaller for Document
type Email ¶
type File ¶
type File struct {
ID uint `json:"id" gorm:"primarykey"`
Provider string `json:"provider"` //IntelX, ...
FilePath string `json:"file_path"`
FileName string `json:"file_name"`
Name string `json:"name"`
Date time.Time `json:"date"`
Bucket string `json:"bucket"`
MediaType string `json:"media_type"`
IndexedAt time.Time `json:"indexed_at"`
Size uint `json:"size"`
ProviderId string `json:"provider_id"`
MIMEType string `json:"mime_type"`
Fingerprint string `json:"fingerprint" gorm:"unique;not null"`
Content string `json:"content"`
// Failed flag set if the result should be considered failed
Failed bool `json:"failed"`
FailedReason string `json:"failed_reason"`
Credentials []Credential `json:"credentials" gorm:"constraint:OnDelete:CASCADE"`
Emails []Email `json:"emails" gorm:"constraint:OnDelete:CASCADE"`
URLs []URL `json:"urls" gorm:"constraint:OnDelete:CASCADE"`
Phones []Phone `json:"phones" gorm:"constraint:OnDelete:CASCADE"`
Documents []Document `json:"documents" gorm:"constraint:OnDelete:CASCADE"`
}
Name,Date,Bucket,Media,Content Type,Size,System ID
func (*File) BeforeCreate ¶ added in v0.1.42
BeforeCreate targets the upsert at the fingerprint instead of the primary key. The connection-level OnConflict set by DbWriter has no Columns, so gorm resolves it to ON CONFLICT ("id") — which never fires for a new row and lets the insert hit uni_files_fingerprint instead. The fingerprint is the real identity of a file here (the Elasticsearch writer uses it as the document _id), so a re-import of the same content must update the existing row.
Declared on the model rather than on the connection so the cascaded inserts into credentials/urls/emails/phones/documents keep their own conflict target; those tables have no fingerprint column.
type Finding ¶
type Finding struct {
// Rule is the name of the rule that was matched
RuleID string
Description string
StartLine int
EndLine int
StartColumn int
EndColumn int
Line string `json:"-"`
Match string
// Secret contains the full content of what is matched in
// the tree-sitter query.
Secret string
// File is the name of the file containing the finding
File string
SymlinkFile string
Commit string
Link string `json:",omitempty"`
// Entropy is the shannon entropy of Value
Entropy float32
Author string
Date string
Message string
Tags []string
// unique identifier
Fingerprint string
Credential Credential
Email Email
Url URL
Phone Phone
Document Document
}
Finding contains information about strings that have been captured by a tree-sitter query.
type LeakIndexable ¶ added in v0.1.39
type LeakIndexable interface {
LeakID() string
LeakType() string
LeakDoc() map[string]interface{}
RefDoc() map[string]interface{}
}
LeakIndexable is implemented by every leak type (Credential, URL, Email, Phone, Document). It powers the restructured Elasticsearch writer, which splits each leak into three concerns:
- LeakID: a content-only fingerprint used as the _id in the global, deduplicated leak index. It must NOT depend on the file it was found in nor on a timestamp, so the same leak seen in different files/imports collapses to a single document.
- LeakDoc: the intrinsic leak fields (the value itself). This is all that gets stored in the leak index — no file reference.
- RefDoc: the occurrence-specific context (near_text, line, source...) that belongs to the monthly file<->leak reference index, not to the leak itself.
- LeakType: a discriminator string for the reference index.
type Phone ¶ added in v0.1.36
type Phone struct {
ID uint `json:"id" gorm:"primarykey"`
FileID uint `json:"file_id" gorm:"index:idx_phone"`
Time time.Time `json:"time"`
Country string `json:"country"`
Raw string `json:"raw"`
Phone string `json:"phone"`
Source string `json:"source"`
FileName string `json:"file_name"`
Line string `json:"line"`
NearText string `json:"near_text"`
}
func (Phone) MarshalJSON ¶ added in v0.1.36
Custom Marshaller for Phone