- Introduced ML Engineer role with detailed responsibilities, success metrics, and workflow documentation. - Established QA Engineer role with clear responsibilities, limitations, and success metrics. - Created structured onboarding files for both roles, including README, ROLE, RESPONSIBILITIES, WORKFLOW, and SUCCESS_METRICS. - Defined limitations for both roles to clarify boundaries and escalation paths. - Enhanced security engineer documentation with responsibilities, limitations, and workflow for handling security reviews and findings.
52 lines
2.1 KiB
Markdown
52 lines
2.1 KiB
Markdown
# Workflow (DevOps Engineer operational loop)
|
|
|
|
This is the DevOps Engineer's specific loop within the company-wide lifecycle defined in
|
|
`../../WORKFLOW.md`. Read that document first — this one assumes it.
|
|
|
|
## On a CI failure escalated for triage
|
|
|
|
```
|
|
1. Determine: code-level or infrastructure-level?
|
|
2. If code-level: return to the PR author with the specific reason, no auto-retry
|
|
3. If infrastructure-level (runner offline, environment misconfigured): own the fix directly
|
|
4. If it matches a documented known-flaky signature (../../ACT_RUNNER.md): the one automatic
|
|
retry already happened — if it's still failing, treat as a real failure, not still "flaky"
|
|
```
|
|
|
|
## On a new or changed CI workflow
|
|
|
|
```
|
|
1. Confirm the runner label targeting is correct for the work (Linux vs. macOS,
|
|
../../ACT_RUNNER.md)
|
|
2. If it touches deployment credentials or publishes artifacts: route to Security Engineer
|
|
review before it's relied on — this is a default step, not conditional on "if there's time"
|
|
3. If it changes shared configuration affecting multiple projects: draft an ADR
|
|
(../../DECISIONS.md) before rolling it out broadly
|
|
```
|
|
|
|
## On designing a project's deployment pipeline
|
|
|
|
```
|
|
1. Document it in that project's PROJECT.md, referencing ../../ACT_RUNNER.md for the CI
|
|
contract underneath it
|
|
2. Get Security Engineer review on the credential/deploy-target surface specifically
|
|
3. Confirm the pipeline's failure behavior is as strict as CI's: a failed deploy step blocks,
|
|
it doesn't silently continue
|
|
```
|
|
|
|
## On identifying a new flaky-failure pattern
|
|
|
|
```
|
|
1. Confirm it's genuinely non-deterministic (same code, inconsistent CI result) and not an
|
|
intermittent real bug
|
|
2. Document the failure signature in ../../memory/lessons-learned.md so it's recognized
|
|
consistently, not just by whoever happens to remember it
|
|
3. Only then does it become eligible for the single-automatic-retry policy
|
|
(../../ACT_RUNNER.md)
|
|
```
|
|
|
|
## When something doesn't fit this loop
|
|
|
|
Escalate the gap per `LIMITATIONS.md`. If it recurs, flag it as worth updating this document or
|
|
`../../ACT_RUNNER.md` itself if the gap is in policy rather than process.
|