# SharePoint Archive Worker

This Node.js 22 service owns the complete archival workflow. Salesforce is passive: it exposes standard REST APIs and stores final `SharePoint_Archived_File__c` audit records only.

The worker:

1. Queries eligible Salesforce Files and legacy Attachments.
2. Resolves supported linked parents and EmailMessage relationships.
3. Streams the source binary to restricted Ubuntu temporary storage.
4. Uploads it to SharePoint with Microsoft Graph.
5. Upserts a final Salesforce audit row using `Archive_Key__c`.
6. Revalidates links/version state and deletes the Salesforce source when deletion is enabled.

If upload or audit creation fails, the source is retained and retry timing is stored in the local failure-state file. If audit creation succeeds but deletion fails, the next scan sees the successful audit and retries deletion without uploading again.

## Build and test

```bash
npm install
npm test
```

Production requires Node.js 22 LTS. There are no runtime npm dependencies.

## Salesforce identity

Create a Connected App that permits JWT bearer authentication and assign the `SharePoint Archive Integration` permission set to its dedicated integration user.

The user's profile/sharing must additionally provide:

- Read access to eligible `ContentVersion`, `ContentDocumentLink`, `Attachment`, and `EmailMessage` records.
- Blob download access for `ContentVersion.VersionData` and `Attachment.Body`.
- Delete access to `ContentDocument` and `Attachment` when `DELETE_AFTER_ARCHIVE=true`.
- Access to every supported linked parent object so discovery results are complete. Consider the Salesforce `Query All Files` permission if the normal sharing model does not expose every archived File.

No custom Apex API, queue, batch, or callout is used.

## Microsoft identity

Register an Entra application with a client secret. Grant the Graph `Sites.Selected` application permission, grant admin consent, and then grant that application `write` permission on the target SharePoint site. Store the secret value only in the protected runtime environment file.

## Ubuntu deployment

1. Build with Node.js 22 and copy `dist`, `package.json`, and the generated lockfile to `/opt/sharepoint-archive-worker`.
2. Create a non-login `sharepoint-archive` user and `/var/lib/sharepoint-archive-worker/tmp` owned by that user.
3. Store private keys under `/etc/sharepoint-archive` with mode `0600` and create `/etc/sharepoint-archive-worker.env` from `.env.example`.
4. Install `sharepoint-archive-worker.service` under `/etc/systemd/system`, then enable and start it.
5. Start with `DELETE_AFTER_ARCHIVE=false`. After reconciling pilot byte counts, URLs, and audit records, change it to `true` and restart the service.

The worker writes structured JSON logs to stdout for collection by journald. Durable retry state is stored at `FAILURE_STATE_FILE`.

## Salesforce deployment order

1. Deploy `manifest/sharepoint-archive-worker.xml`. This makes the legacy scheduler a safe no-op and adds the final audit fields/permission set.
2. Confirm no legacy batch is running.
3. Apply `manifest/sharepoint-archive-worker-destructive.xml` to remove the former Apex archive batches, callout service, queue API, and queue-only fields if they were previously deployed.
4. Do not enable source deletion until the Node worker pilot succeeds.
