Modelo de benchmarks
dispatchatlas.bench define a evidência de benchmarks antes de os solvers ou as campanhas
a consumirem. Uma família de benchmarks declara sua taxonomia, perfil de domínio, classe
de perfil, suposições, evidência de citação, envelope de escala, espaço de nomes de
semente, e esquema de saída. A materialização valida cada problema gerado com
dispatchatlas.core, caracteriza a instância, envolve-a em um envelope de proveniência, e
registra hashes estáveis.
O catálogo cobre famílias genéricas de escalonamento de otimização-combinatória e escalonamento de distributed-computing como pares co-iguais, de modo que a plataforma não é uma ferramenta apenas-de-distributed-computing.
Exemplo executável: examples/benchmark_continuum.py gera, caracteriza, e cataloga um atlas de benchmarks contínuo em tempo real.
Famílias de escalonamento
Cada família de escalonamento é materializada como um par de catálogo de primeira-classe com ao menos um perfil gerador. A tabela de catálogo abaixo é gerada a partir do registro de geradores de benchmarks e da matriz de citação, de modo que seus totais de família e citações de fonte são contáveis a partir das próprias linhas. Um gráfico de distribuição de famílias acima da tabela mostra como os perfis de família se distribuem entre as categorias de famílias de escalonamento.
Generated from the benchmark generator registry and the citation matrix: 69 family profiles across 8 scheduling families — distributed-computing (47), flow-shop (7), job-shop (6), machine-scheduling (3), open-shop (1), rcpsp (3), rcpsp-max (1), setup-flow-shop (1).
Showing 69 of 69 family profiles.
| Distinctive against | Evidence | Citation status | Source citations | |||
|---|---|---|---|---|---|---|
accelerator-coschedulingaccelerator-coscheduling a heterogeneous datacenter job needs a compute resource and a scarce accelerator at the same time, so every job co-allocates two simultaneous resource demands held together for its whole run; the accelerator pool is scarce, so jobs sharing an accelerator serialize on it while jobs on disjoint resources run in parallel, and the scheduler reasons over a multi-resource co-allocation problem rather than a single-unit-demand one | distributed-computing | ioe-complete | published single-resource continuum schedulers (one unit resource demand per task, not the simultaneous co-allocation of a compute resource and a scarce accelerator held together under a Pareto contract) | smoke | citation-backed |
|
aerial-edgeaerial-edge a loitering UAV is a flying fog node that serves the ground region beneath it for a fixed loiter window before moving to the next pass, so sorties group into successive loiter passes pinned to the fog tier the platform embodies while overhead | distributed-computing | ioe-complete | aerial-edge MEC simulators (no flying-fog loiter placement) | smoke | citation-backed |
|
anytime-inferenceanytime-inference an edge accelerator serves deep-learning inference requests that each complete a mandatory minimal-accuracy early-exit branch and may run an optional refinement to full accuracy when their latency deadline allows; every request declares a mandatory duration below its full duration and a tight latency deadline, and arrivals are spaced shorter than a full inference so requests queue and contend, so the scheduler decides which requests refine and which deliver the early-exit result -- an imprecise-computation quality-versus-timeliness trade-off | distributed-computing | ioe-complete | published edge-cloud split inference and datacenter inference serving (a fixed full-accuracy computation per request), neither of which lets a request drop an optional refinement to meet its deadline so the schedule order trades accuracy for timeliness | smoke | citation-backed |
|
blocking-flow-shopblocking-flow-shop | flow-shop | classical | — | smoke | citation-backed |
|
bulk-synchronous-graphbulk-synchronous-graph an iterative graph computation runs as a sequence of supersteps separated by global barriers, so each superstep's vertex partitions compute and exchange messages and every partition of the next superstep waits on all partitions of the prior one; the slowest partition therefore gates each superstep, and the scheduler reasons over a barrier-synchronized partition-balancing problem rather than an overlap-friendly pipeline or an independent-task one | distributed-computing | ioe-complete | published pipeline or independent-task schedulers (an overlap-friendly wavefront or unsynchronized tasks, not supersteps separated by global barriers where the slowest partition gates each round under a Pareto contract) | smoke | citation-backed |
|
carbon-awarecarbon-aware flexible jobs defer to low-carbon-intensity windows under a time-varying grid carbon signal while honoring their SLA deadlines | distributed-computing | ioe-complete | Electricity-Maps grid carbon-intensity and CityLearn carbon-aware community signals | smoke | citation-backed |
|
cloud-independentcloud-edge-independent | distributed-computing | domain-specific | — | smoke | citation-backed |
|
coflow-schedulingcoflow-scheduling a distributed-computing stage completes only when the last network transfer of its coflow lands, not the first, so a coflow's completion time is the maximum over its member flows; every coflow emits data-heavy flow tasks that place freely across the fabric plus a barrier task that depends on all of them, so the barrier gates the group and the coflow-completion-time is an all-or-nothing footprint | distributed-computing | ioe-complete | datacenter coflow schedulers (no continuum tier-placement barrier) | smoke | citation-backed |
|
compact-job-shopcompact-job-shop | job-shop | classical | — | smoke | citation-backed |
|
confidential-edgeconfidential-edge tasks are classified by data sensitivity -- a confidential task that touches protected data must execute inside the enclave-capable trusted tier so its data never leaves the trusted boundary, while a public task draws a free placement affinity across the fabric | distributed-computing | ioe-complete | edge enclave runtimes (no security-classified Pareto placement) | smoke | citation-backed |
|
cyber-physicalcyber-physical control cycles arrive periodically under a time-varying tariff | distributed-computing | ioe-complete | periodic hard-real-time task models and smart-grid demand-side scheduling formulations (no edge-fog-cloud tier placement under a Pareto contract) | smoke | citation-backed |
|
data-localitydata-locality a query over a large dataset is cheaper to run where the data already resides than to ship the data across the WAN fabric; each task's input lives on one tier (edge sensor logs, fog warm aggregates, or cloud cold archives) and the task pins to that tier so its heavy input never crosses the fabric, with the data tiers cycled so placement spans the whole edge-fog-cloud continuum | distributed-computing | ioe-complete | cluster locality schedulers (no continuum data-residency placement) | smoke | citation-backed |
|
datacenter-colocationdatacenter-colocation latency-sensitive service jobs and deferrable batch jobs share multi-tenant cells, so high-priority-band jobs claim capacity ahead of low-band jobs under contention | distributed-computing | ioe-complete | Google Borg ClusterData2019 priority-tiered cell traces | smoke | citation-backed |
|
digital-twin-syncdigital-twin-sync each physical asset periodically syncs its state to its fog/cloud twin and must finish within a freshness (Age-of-Information) window before the twin's state goes stale | distributed-computing | ioe-complete | digital-twin edge frameworks (no joint Pareto placement) | smoke | citation-backed |
|
disaggregated-memorydisaggregated-memory a CXL-pooled cloud platform backs each socket with a small local DRAM tier and a shared far-memory pool, and every VM draws a long-tailed memory working set, so a few memory-hungry tenants dominate a socket's local budget while far-memory access inflates a VM's runtime in proportion to the working set it spills to the pool; the scheduler reasons over local-versus-pool placement rather than core-count placement | distributed-computing | ioe-complete | published CXL memory-pooling and tiered-memory systems (socket-local page placement, no edge-fog-cloud tier scheduling under a Pareto contract) | smoke | citation-backed |
|
distributed-assembly-flow-shopdistributed-assembly-flow-shop | flow-shop | classical | — | smoke | citation-backed |
|
distributed-flexible-job-shopdistributed-flexible-job-shop | job-shop | structurally-complex | — | smoke | citation-backed |
|
distributed-permutation-flow-shopdistributed-permutation-flow-shop | flow-shop | classical | — | smoke | citation-backed |
|
distributed-training-gangdistributed-training-gang a GPU cluster runs synchronous data-parallel training jobs; each job is a gang of workers that must START TOGETHER on distinct accelerators (every all-reduce step synchronizes the workers), so a job cannot begin until enough accelerators are free simultaneously; the workers reuse the accelerator pool and jobs arrive over time, so jobs queue and the scheduler decides which job acquires a full simultaneously-free worker set first -- an all-or-nothing gang co-start, not an independent placement of each worker | distributed-computing | ioe-complete | published GPU-cluster and inference-serving families that place each task independently; none requires a whole job's worker set to co-start simultaneously on distinct accelerators, so no other family forbids a partial start -- the gang-scheduling all-or-nothing constraint under a Pareto contract | smoke | citation-backed |
|
distributed-transactiondistributed-transaction a partitioned database runs transactions that each acquire exclusive locks on a variable read/write set of data shards, so every transaction co-allocates a randomly drawn subset of shards held together for its whole run; two transactions whose shard sets intersect serialize while disjoint transactions commit in parallel, so the scheduler reasons over a variable-cardinality lock-conflict graph rather than a fixed two-resource hold | distributed-computing | ioe-complete | published replica-placement schedulers (a single shard pinned per task for locality, not a variable-cardinality exclusive lock set co-allocated per transaction forming a conflict graph under a Pareto contract) | smoke | citation-backed |
|
edge-offloadingedge-offloading-mec each task chooses between local edge execution and remote offload | distributed-computing | ioe-complete | iFogSim MEC offloading scenarios | smoke | citation-backed |
|
edge-placementedge-placement services place on edge servers near their user population and migrate as demand shifts across base-station coverage cells | distributed-computing | ioe-complete | EUA edge-user-allocation and Shanghai-Telecom base-station placement traces | smoke | citation-backed |
|
elastic-serverless-autoscaleelastic-serverless-autoscale a serverless platform serves function invocations on a shared worker pool, and every invocation is moldable: it may run on one, two, or four concurrent workers, where a wider allocation runs shorter by a sublinear speedup but spends more total worker-seconds; each request declares its execution modes and arrivals are spaced shorter than a base invocation so the pool is contended, so the scheduler picks each invocation's worker width -- a moldable latency-versus-cost choice rather than a fixed resource hold | distributed-computing | ioe-complete | published serverless cold-start and fixed-width co-allocation families (accelerator co-scheduling, fpga partitioning), each of which holds one fixed resource set per task; none lets a request choose among several worker-count modes so the schedule order trades latency for resource cost under a Pareto contract | smoke | citation-backed |
|
facility-assignmentfacility-assignment | machine-scheduling | classical | — | smoke | citation-backed |
|
failure-recoveryfailure-recovery a failed task re-places its checkpoint state to a surviving tier | distributed-computing | ioe-complete | Borg cluster failure-event traces | smoke | citation-backed |
|
federated-learningfederated-learning each training round selects a subset of heterogeneous, straggler-prone edge clients that train on non-IID local data, then a fog or cloud aggregator combines their updates | distributed-computing | ioe-complete | FedScale and Oort federated-learning device-participation benchmarks | smoke | citation-backed |
|
flexible-job-shopflexible-job-shop | job-shop | classical | — | smoke | citation-backed |
|
flow-shoppermutation-flow-shop | flow-shop | classical | — | smoke | citation-backed |
|
fpga-partitioningfpga-partitioning a multi-tenant reconfigurable FPGA hosts tenant kernels that each occupy a contiguous region of fabric tiles, so every kernel co-allocates a contiguous run of tiles held together for its whole residency; two kernels whose tile intervals overlap cannot be co-resident and serialize while kernels on disjoint tile spans run in parallel, so the scheduler reasons over an interval-overlap conflict graph rather than a fixed two-resource hold or a random-subset lock set | distributed-computing | ioe-complete | published accelerator co-scheduling (a fixed compute-plus-accelerator pair) and shard-lock transactions (a random subset of resources), neither of which constrains the co-allocated set to a spatially contiguous tile interval whose overlaps form an interval conflict graph under a Pareto contract | smoke | citation-backed |
|
frontierco-fjspfrontierco-fjsp | job-shop | classical | — | smoke | citation-backed |
|
generative-inference-servinggenerative-inference-serving a transformer inference replica batches autoregressive requests that each hold key-value-cache memory proportional to their token count for the whole decode, so a long-tailed sequence mix fragments a fixed cache budget and the scheduler reasons over memory-bound admission rather than GPU-count placement; every request's decode duration and KV-cache demand scale with its drawn token count | distributed-computing | ioe-complete | published single-replica generative-model serving systems (no edge-fog-cloud tier placement under a Pareto contract) | smoke | citation-backed |
|
gpu-mlgpu-ml training and inference jobs claim accelerators and gang-schedule replicas | distributed-computing | ioe-complete | Alibaba PAI, Philly, and Helios GPU-cluster traces | smoke | citation-backed |
|
hybrid-flow-shophybrid-flow-shop | flow-shop | classical | — | smoke | citation-backed |
|
immersive-xrimmersive-xr each extended-reality frame runs a latency-critical perception, render, and display pipeline placed across the device, edge, and cloud within a hard motion-to-photon deadline | distributed-computing | ioe-complete | ILLIXR extended-reality systems testbed | smoke | citation-backed |
|
intermittent-edgeintermittent-edge a batteryless sensor harvests ambient energy into a small buffer, runs until the buffer depletes, then sleeps to recharge; a job too large for one duty-cycle window is checkpointed at power loss and resumed in the next, so it is a precedence chain of edge-pinned per-window segments each bounded by the constant energy window | distributed-computing | ioe-complete | intermittent-computing runtimes (no continuum energy-window placement) | smoke | citation-backed |
|
iot-edgeiot-edge many small sensor readings arrive periodically and aggregate at the edge | distributed-computing | ioe-complete | published wireless-sensor-network telemetry datasets and in-network aggregation deployments (raw sensor readings, not edge-fog-cloud tier scheduling under a Pareto contract) | smoke | citation-backed |
|
job-shopjob-shop | job-shop | classical | — | smoke | citation-backed |
|
kv-cache-placementkv-cache-placement a distributed key-value cache tier serves a catalog of cache objects whose request rate follows a heavy Zipfian popularity skew, so a few hot objects absorb most of the traffic; each object is placement-flexible, carrying one single-node mode per cache-tier node, so the scheduler chooses which tier node hosts it, and an object's working set -- the transfer volume staged onto its host tier -- scales with its popularity rank, so the hottest object carries the largest working set and a cold-tail object the smallest; a naive uniform placement strands a hot, large-working-set object on a far tier and pays its transfer across the fabric, while a locality-aware placement pins the hottest objects to near tiers to shrink makespan and cost | distributed-computing | ioe-complete | published cache and content-placement families (consistent-hashing replica placement, CDN content distribution) that place each object uniformly or by a hash; none scales each object's working set with a Zipfian popularity rank so the hot objects carry a strictly larger transfer volume, making popularity-skewed near-tier pinning the lever a locality-aware placement pulls under the Pareto contract | smoke | citation-backed |
|
machine-schedulingmachine-scheduling-unrelated | machine-scheduling | classical | — | smoke | citation-backed |
|
microservice-dagmicroservice-dag services form an acyclic call graph pinned by role to a tier | distributed-computing | ioe-complete | Alibaba v2021 microservice-trace call graphs | smoke | citation-backed |
|
mixed-criticalitymixed-criticality a safety-critical real-time mix runs tasks of differing criticality, and a high-criticality task is budgeted with a conservative high-assurance worst-case execution time and a tight deadline while a low-criticality task carries a smaller best-effort budget and a loose deadline, so the criticality tiering lives in the duration and deadline structure; the scheduler reasons over which assured-criticality tasks to guarantee under contention rather than a uniform-assurance deadline-scheduling one | distributed-computing | ioe-complete | published uniform-assurance real-time deadline schedulers (one worst-case execution time and deadline class per task, not criticality-tiered WCET budgets with tighter high-assurance deadlines under a Pareto contract) | smoke | citation-backed |
|
moe-expert-parallelmoe-expert-parallel a sparsely-activated mixture-of-experts model routes each token batch to one expert and the experts are spread across devices, and expert popularity is long-tailed, so a few hot experts receive most token batches while many stay cold and the all-to-all routing exchange dominates fabric traffic; the scheduler reasons over an expert-placement and load-balancing problem rather than a dense uniform-replica serving one | distributed-computing | ioe-complete | published dense generative-model serving systems (uniform per-replica KV-cache admission, not sparse token-to-expert routing under load imbalance and a Pareto contract) | smoke | citation-backed |
|
multi-objective-pfspmulti-objective-pfsp | flow-shop | classical | — | smoke | citation-backed |
|
multi-project-rcpspmulti-project-rcpsp | rcpsp | structurally-complex | — | smoke | citation-backed |
|
multi-tenant-fair-sharemulti-tenant-fair-share a shared cluster serves several tenants whose workloads compete for one node pool; each task is placement-flexible, carrying one single-node mode per pool node, so the scheduler chooses which node it occupies; tenants are sized asymmetrically, so even a load-balanced placement leaves the heavy tenants holding a larger fraction of their busiest node -- a higher dominant resource share -- than the light ones, and a fairness-aware scheduler rebalances placement to shrink the dominant-share spread | distributed-computing | ioe-complete | published multi-tenant colocation families (Borg-style priority colocation, vm allocation) that fix each task's resource and score makespan or cost; none lets the scheduler choose each tenant task's node and scores the dominant-resource-share spread between tenants as a fairness objective | smoke | citation-backed |
|
network-slicingnetwork-slicing isolated slice classes (latency-critical, broadband, massive-IoT) each carry their own service-level deadline and placement, and same-class slices spread across tiers for resilience | distributed-computing | ioe-complete | 5G slicing orchestration (no joint Pareto placement) | smoke | citation-backed |
|
no-wait-flow-shopno-wait-flow-shop | flow-shop | classical | — | smoke | citation-backed |
|
open-shopopen-shop | open-shop | classical | — | smoke | citation-backed |
|
orbital-edgeorbital-edge tasks schedule across ground terminals, moving low-earth-orbit satellites, and cloud backhaul under time-varying connectivity as satellites enter and leave coverage and hand work over | distributed-computing | ioe-complete | LENS real-measurement LEO satellite-network traces | smoke | citation-backed |
|
pipeline-parallel-trainingpipeline-parallel-training a deep network is split into successive pipeline stages pinned across edge-to-cloud tiers and the training mini-batch is divided into micro-batches, so each micro-batch flows forward stage by stage while each stage runs its micro-batches in issue order; the two precedence families form a diagonal wavefront whose warm-up and cool-down idle slots are the pipeline bubbles, and deeper stages carry rising compute, so the scheduler reasons over a stage-partition and bubble-minimizing problem rather than a synchronous data-parallel all-reduce one | distributed-computing | ioe-complete | published data-parallel / gang-scheduled training systems (synchronous all-reduce over co-located replicas, not a stage-by-micro-batch pipeline wavefront with warm-up and cool-down bubbles under a Pareto contract) | smoke | citation-backed |
|
rcpsprcpsp-renewable | rcpsp | structurally-complex | — | smoke | citation-backed |
|
rcpsp-maxrcpsp-max | rcpsp-max | structurally-complex | — | smoke | citation-backed |
|
rcpsp-multi-modercpsp-multi-mode | rcpsp | structurally-complex | — | smoke | citation-backed |
|
reentrant-fabreentrant-fab | job-shop | structurally-complex | — | smoke | citation-backed |
|
replica-placementreplica-placement a replica runs where its data shard already lives | distributed-computing | ioe-complete | CRUSH replicated-data placement | smoke | citation-backed |
|
serverless-cold-startserverless-cold-start a cold invocation pays a container provisioning penalty | distributed-computing | ioe-complete | Azure Functions serverless-in-the-wild traces | smoke | citation-backed |
|
service-function-chainservice-function-chain an NFV packet flow traverses a linear ordered chain of typed virtual network functions -- firewall, intrusion detection, deep packet inspection, address translation -- each pinned to a tier that hosts its function type, so the chain is a strict total order and the flow crosses the edge-fog-cloud fabric in a fixed sequence; the scheduler reasons over a chain-placement problem under an end-to-end latency budget rather than the branching role-pinned call graph of a microservice | distributed-computing | ioe-complete | published microservice call-graph schedulers (a branching role-pinned acyclic call graph, not a strict linear chain of function-typed network functions under an end-to-end latency budget and a Pareto contract) | smoke | citation-backed |
|
setup-flow-shopsetup-flow-shop | setup-flow-shop | classical | — | smoke | citation-backed |
|
smartnic-offloadsmartnic-offload a SmartNIC-accelerated server pairs a fast host CPU with a low-power on-NIC processor, and every microservice draws a long-tailed compute intensity, so most are light enough to offload onto the energy-frugal NIC cores while a few compute-heavy services must stay host-bound; a service's runtime scales with its intensity, so the scheduler reasons over an energy-versus-latency offload-placement problem rather than a uniform host placement one | distributed-computing | ioe-complete | published mobile-edge computation-offloading models (device-to-edge latency offload, not in-server host-to-NIC energy offload under a Pareto contract) | smoke | citation-backed |
|
split-inference-servingsplit-inference-serving each inference request partitions a deep model at a layer cut -- a light head runs the early layers on the edge near the sensor and a heavy tail runs the later layers in the cloud, consuming the head's intermediate feature map under a per-request end-to-end latency SLO | distributed-computing | ioe-complete | datacenter inference serving (no edge-cloud partition placement) | smoke | citation-backed |
|
spot-preemptiblespot-preemptible a cloud provider rents idle capacity at a discount as revocable spot instances reclaimed after a short lease; eviction-tolerant batch work pins to the cloud spot tier under a hard lease deadline (the eviction horizon), while latency-critical interactive work pins to the stable edge on-demand tier with no eviction deadline | distributed-computing | ioe-complete | cloud spot schedulers (no continuum eviction-deadline placement) | smoke | citation-backed |
|
storage-io-tieringstorage-io-tiering a tiered storage pool serves I/O-bound jobs whose runtime is dominated by moving a job's I/O volume through the storage node it lands on; each job is placement-flexible, carrying one single-node mode per tier node, and the mode duration is tier-dependent -- a seek floor plus the I/O volume divided by that tier's I/O bandwidth, which differs by tier (a fast cloud array sustains far more bytes/second than a slow edge disk); a job's I/O volume follows a heavy-tailed falloff over its I/O-demand rank, so a few I/O-heavy jobs carry most of the bytes and have a large cross-tier duration spread, while the light tail barely varies; a naive placement strands an I/O-heavy job on a low-bandwidth tier and pays its volume slowly, while a bandwidth-aware placement pins the heavy jobs to fast tiers to shrink makespan and cost | distributed-computing | ioe-complete | published storage-tiering and hierarchical-storage-management families that migrate blocks between fast and slow tiers by access frequency; none models the heterogeneous-tier I/O bandwidth as a placement-flexible per-tier mode whose duration is the I/O volume divided by that tier's bandwidth, so an I/O-heavy job's cross-tier duration spread is the bottleneck-relief lever a bandwidth-aware placement pulls under the Pareto contract | smoke | citation-backed |
|
streaming-windowstreaming-window events arrive online in bounded windows and must close within one | distributed-computing | ioe-complete | Parallel Workloads Archive online arrivals | smoke | citation-backed |
|
time-sensitive-networkingtime-sensitive-networking each time-triggered flow releases on a fixed period and must finish within one cycle under a hard, jitter-free deadline | distributed-computing | ioe-complete | EdgeCloudSim best-effort scenarios (no gating) | smoke | citation-backed |
|
unrelated-parallel-setupunrelated-parallel-setup | machine-scheduling | structurally-complex | — | smoke | citation-backed |
|
vehicular-offloadingvehicular-offloading a vehicle's tasks share an arrival time and a roadside-unit dwell deadline, and hand over from the roadside unit to the fog tier as the vehicle drives on | distributed-computing | ioe-complete | EdgeCloudSim / SUMO vehicular-edge mobility scenarios | smoke | citation-backed |
|
video-analyticsvideo-analytics each camera streams frames that must be analyzed within a tight real-time latency bound, placed hierarchically with edge inference near the camera and cloud aggregation | distributed-computing | ioe-complete | edge video-analytics clusters (no joint Pareto placement) | smoke | citation-backed |
|
vm-allocationvm-allocation size-heterogeneous virtual-machine deployments pack onto hosts while each deployment's members spread across distinct failure domains for availability | distributed-computing | ioe-complete | Azure Public Dataset Resource Central VM-allocation traces | smoke | citation-backed |
|
workflow-dagworkflow-dag | distributed-computing | structurally-complex | — | smoke | citation-backed |
|
Per-instance characterization and download eligibility live in the benchmark catalog; this table is the family-and-citation inventory.
O componente acima carrega o inventário gerado completo — as famílias núcleo clássicas e de distributed-computing abaixo mais o contínuo de famílias Edge–Fog–Cloud. Essas famílias núcleo fundacionais em forma de documento:
| Família | Perfil | Classe de perfil | Corpus primário |
|---|---|---|---|
| Escalonamento de máquinas (R||Cmax, máquina não-relacionada) | machine-scheduling-unrelated | classical | OR-Library |
| Job-shop | job-shop-classical | classical | OR-Library, Taillard |
| Job-shop flexível (FJSP) | flexible-job-shop | classical | Brandimarte; Hurink-Jurisch-Thole |
| Flow-shop de permutação | permutation-flow-shop | classical | Taillard |
| Flow-shop de setup dependente-de-sequência (SDST) | setup-flow-shop | classical | Allahverdi et al. (2008); Allahverdi (2015) |
| Escalonamento de projetos com-recursos-restritos (RCPSP) | rcpsp-renewable | structurally-complex | PSPLIB |
| Distributed-computing (cloud/edge) | cloud-edge-capacity | domain-specific | CloudSim; DynamicCloudSim; Edge vision |
| Distributed-computing (workflow DAG) | workflow-dag | structurally-complex | Standard Task Graph Set |
As famílias de máquina-não-relacionada e job-shop flexível anexam uma matriz de
tempo-de-execução CostModel a cada instância, exportada como uma matriz materializada
junto ao JSON do problema. A família de flow-shop de setup dependente-de-sequência em vez
disso anexa uma matriz de setup CostModel: uma troca entre trabalhos de famílias
diferentes em uma máquina custa tempo de setup, de modo que o objetivo de setup recompensa
agrupar trabalhos similares. Os corpora padrão nomeados são citados e vinculados apenas e
nunca são redistribuídos dentro do repositório.
Várias famílias do contínuo exercitam co-alocação multi-recurso: cada tarefa demanda
mais de um recurso ao mesmo tempo e o construtor os mantém juntos por toda a sua duração
(ver Contratos de domínio). A família accelerator-coscheduling
co-aloca um nó de computação e um acelerador escasso por trabalho; a família
distributed-transaction co-aloca um conjunto de bloqueios de cardinalidade-variável de
fragmentos de dados por transação; e a família fpga-partitioning co-aloca uma corrida
espacialmente contígua de tiles de tecido reconfigurável por kernel de inquilino. As
tarefas cujos conjuntos de recursos se intersectam serializam enquanto tarefas disjuntas
rodam concorrentemente — o executável
examples/inspect_coallocation.py
torna a alavanca explícita.
Além da co-alocação, três famílias do contínuo exercitam suas próprias alavancas
estruturais. A família elastic-serverless-autoscale exercita execução moldável: cada
invocação de função declara mais de um modo de execução — um modo estreito
apenas-em-casa e um modo largo que empresta um worker de um pequeno pool de rajada
compartilhado para terminar mais cedo — de modo que o escalonamento escolhe um modo por
tarefa e a ordem decide quais invocações reclamam o escasso modo largo-e-rápido. A família
distributed-training-gang exercita co-escalonamento de gangue: os workers de um
trabalho de treinamento data-paralelo síncrono compartilham uma gangue e devem co-iniciar
em aceleradores distintos em um lançamento tudo-ou-nada — os workers reutilizam o pool de
aceleradores e os trabalhos chegam ao longo do tempo, de modo que um trabalho não pode
começar até que aceleradores suficientes se liberem simultaneamente, e a ordem decide qual
trabalho adquire seu conjunto completo de workers primeiro. A família
multi-tenant-fair-share exercita equidade de recurso-dominante: vários inquilinos de
tamanho-assimétrico colocam tarefas flexíveis-em-colocação em um pool de nós compartilhado,
e o objetivo de cota-de-recurso-dominante pontua a dispersão entre a cota dominante do
inquilino mais- e menos-servido — de modo que a colocação, quais recursos cada inquilino
ocupa, é a alavanca que a equilibra ou a enviesa. Os executáveis
examples/serverless_autoscale_study.py,
examples/distributed_training_gang_study.py,
e
examples/multi_tenant_fairshare_study.py
tornam essas três alavancas explícitas.
Classes de perfil
| Classe de perfil | Significado |
|---|---|
classical | Deriva de um corpus padrão de otimização-combinatória. |
structurally-complex | Carrega estrutura de precedência, DAG, ou rede-de-recursos. |
ioe-complete | Cenário distribuído Internet-of-Everything-completo. |
trace-backed | Fundamentado em um traço de carga do mundo-real nomeado. |
domain-specific | Adaptado a um único domínio operacional. |
Rótulos de evidência
| Rótulo | Uso |
|---|---|
smoke | Instâncias determinísticas pequenas para testes, exemplos, docs, e prévias. |
exploratory | Material plausível que ainda não é respaldado-por-citações nem totalmente caracterizado. |
| grau de evidência candidato | Material respaldado-por-citações à espera de gates de piloto, estatísticos, e de campanha. |
| grau de evidência de campanha-completa | Evidência que passou pelos gates de citação, caracterização, estatísticos, de divulgação, e de qualidade. |
Os catálogos de fumaça nunca são evidência de avaliação final. Existem para provar que os
geradores, a validação, a caracterização, as verificações de citação, e a persistência
funcionam rapidamente. O catálogo do portal e os downloads previsualizam cada família do
contínuo nesta pequena escala de fumaça (um pool de três-recursos); uma família de
co-alocação não pode exibir paralelismo de recurso-disjunto em um pool tão pequeno, de modo
que a estrutura distintiva é uma propriedade de escala-de-pesquisa.
build_continuum_full_catalog() materializa cada família em sua escala de pesquisa
declarada -- o pool de recursos e a contagem de tarefas maiores onde a estrutura de
co-alocação, contenção, e colocação genuinamente se manifesta -- para pacotes de benchmarks
de grau-pesquisa.
Taxonomia
A taxonomia cobre estruturas de escalonamento, ambientes, realismo de infraestrutura, características de objetivo, características de restrição, incerteza, e dinamismo. Os exemplos incluem workflows DAG, lotes de tarefas independentes, funções serverless, consolidação de contêineres e VM, ambientes edge e cloud, traços públicos, otimização multi-objetivo, deadlines, localidade de dados, churn, e chegadas dinâmicas.
Matriz de citação
As afirmações de benchmarks são verificadas contra CitationMatrix. As afirmações de grau
de evidência candidato e de campanha-completa falham na validação a menos que referenciem
fontes respaldadas-por-citações. Material não suportado deve permanecer exploratory até que
evidência seja adicionada.
O conjunto de fontes é declarado em default_citation_matrix() com identificadores de
fonte estáveis, referências resolúveis, e uma postura de licença registrada. Abrange três
níveis: os corpora padrão de otimização-combinatória (citados e vinculados apenas, nunca
empacotados), traços de cluster de produção, e um amplo conjunto de datasets contemporâneos
do mundo-real do contínuo Edge–Fog–Cloud — traços de cluster de GPU e machine-learning,
suites de benchmark de microsserviços e serverless, traços de workflow-científico, traços
de trabalhos de supercomputador, datasets de edge-placement e mobilidade, datasets de IoT
e demanda-celular, sinais de carbono e energia de rede, cargas de processamento-de-fluxo,
benchmarks de participação-de-dispositivos de federated-learning, testbeds de sistemas de
realidade-estendida, e traços de rede-satelital de órbita-terrestre-baixa:
Suites de referência canônicas
O registro em default_reference_suites() registra as suites de instâncias publicadas
canônicas às quais cada família de escalonamento genérica se ancora: identidade, família,
contagem de instâncias, ponteiro de recuperação, e os rastreadores de
melhores-soluções-conhecidas que publicam limites para a suite. As suites clássicas são
apenas citadas e vinculadas — o DispatchAtlas nunca empacota nem redistribui arquivos de
instâncias de terceiros.
| Suite | Família | Instâncias | Download | Rastreador BKS |
|---|---|---|---|---|
fisher-thompson | job-shop | 3 | OR-Library | van-hoorn-2018, scheduleopt-benchmarks |
lawrence | job-shop | 40 | Espelho JSPLIB | van-hoorn-2018, scheduleopt-benchmarks |
adams-balas-zawack | job-shop | 5 | Espelho JSPLIB | van-hoorn-2018, scheduleopt-benchmarks |
applegate-cook-orb | job-shop | 10 | Espelho JSPLIB | van-hoorn-2018, scheduleopt-benchmarks |
storer-wu-vaccari | job-shop | 20 | Espelho JSPLIB | van-hoorn-2018, scheduleopt-benchmarks |
yamada-nakano | job-shop | 4 | Espelho JSPLIB | van-hoorn-2018, scheduleopt-benchmarks |
taillard-jsp | job-shop | 80 | Espelho JSPLIB | van-hoorn-2018, scheduleopt-benchmarks |
demirkol-dmu | job-shop | 80 | Espelho JSPLIB | scheduleopt-benchmarks |
brandimarte-mk | job-shop (flexível) | 15 | Espelho SchedulingLab | scheduleopt-benchmarks |
hurink-fjsp | job-shop (flexível) | 198 | Espelho SchedulingLab | scheduleopt-benchmarks |
dauzere-peres-paulli | job-shop (flexível) | 18 | Espelho SchedulingLab | scheduleopt-benchmarks |
taillard-pfsp | flow-shop | 120 | OR-Library | zenodo-pfsp-bks-2021 |
vrf-pfsp | flow-shop | 480 | Site do grupo SOA | zenodo-pfsp-bks-2021 |
sdst-taillard-ruiz | flow-shop de setup | 480 | Site do grupo SOA | as melhores soluções acompanham as instâncias |
cicirello-wt-sds | escalonamento de máquinas | 120 | Harvard Dataverse | cicirello-wtsds-benchmark |
or-library-smtwt | escalonamento de máquinas | 375 | OR-Library | crauwels-potts-vanwassenhove-1998 |
vallada-ruiz-upmsp | escalonamento de máquinas | 1640 (reportado) | Site do grupo SOA | — |
psplib | rcpsp | 2040 | Site do PSPLIB | psplib-1997 |
mmlib | rcpsp | 4320 (reportado) | Página do OR&S | solutionsupdate-ugent-rcpsp |
rg300 | rcpsp | 480 | Página do OR&S | solutionsupdate-ugent-rcpsp |
As suites com um parser empacotado (texto job-shop padrão, matrizes de flow-shop de
Taillard, job-shop flexível .fjs, JSON WfFormat do WfCommons) são ingeridas com
load_reference_suite(suite_id, instances_root=...) a partir de arquivos que o operador
baixa e coloca sob uma árvore resources/ local. A ingestão executa inteiramente offline,
reutiliza o mesmo envelope de validação, caracterização, hashing, e proveniência dos
geradores sintéticos, e carimba cada problema com seu suite_id e upstream_instance_id.
As suites apenas-de-registro são registradas com suas citações e ponteiros de recuperação
sem um parser empacotado.
Registros de melhores-soluções-conhecidas
Os valores de melhor-solução-conhecida por-instância nunca são distribuídos com o
DispatchAtlas. O operador os ingere como arquivos JSON sob um diretório privado
resources/benchmarks/bks/, um arquivo por suite, cada um carregando schema_version, o
suite_id, o source_id do rastreador, a data de recuperação, e as entradas de valor
(id da instância, objetivo, valor, tipo ótimo-ou-limite-superior, limite inferior
opcional). load_best_known_registry valida cada arquivo contra as suites de referência e
a matriz de citação e falha fechando em suites desconhecidas, rastreadores desconhecidos,
entradas duplicadas, ou limites inconsistentes. Sem um registro ingerido, as métricas de
desvio-relativo estão simplesmente indisponíveis — elas nunca são computadas parcialmente,
e nenhum valor de melhor-solução-conhecida aparece em qualquer superfície pública.
Divergências de calibração
As famílias genéricas sintéticas são ancoradas às suites canônicas sem alegar reproduzir seus esquemas de geração. As divergências conhecidas são documentadas em vez de ocultadas:
| Família | Convenção publicada | Convenção sintética |
|---|---|---|
| flow-shop de setup | setups SDST-Taillard a 10/50/100/125% do tempo de processamento | três famílias de setup, custo = família + 1 |
| escalonamento de máquinas (R||Cmax) | classes de duração U[1,100] e variantes de máquina-correlacionada | fatores de velocidade por-par 0.5–2.0 |
As comparações contra as convenções publicadas passam pelas instâncias canônicas ingeridas, não pelas famílias sintéticas.
Caracterização
Cada problema materializado recebe descritores normalizados para:
- densidade de oportunidade
- esparsidade de compatibilidade
- contenção e sobrecarga
- profundidade de dependência
- pressão de comunicação
- intensidade de setup
- viés de carga e heterogeneidade
- conflito de objetivo
- incerteza e dinamismo
- sensibilidade do solver
Ponte de distribution-distance
Cada perfil de grau de evidência de campanha-completa declara uma ponte de distribution-distance:
seu estado (synthetic, calibrated-synthetic, trace-backed, ou externally-sourced),
evidência de calibração, cenário de domínio, cobertura de transferência e disrupção, e
risco de distribution-distance residual. A promoção ao grau de evidência de campanha-completa falha
fechando a menos que as coberturas de transferência e disrupção sejam declaradas, e um
perfil calibrated-synthetic deve nomear a referência trace-backed contra a qual calibra.
As famílias calibrated-synthetic nomeiam sua referência de traço explicitamente; as suites
canônicas ingeridas carregam uma ponte externally-sourced, e o adaptador WfCommons é a
primeira fonte de instâncias trace-backed analisada-externamente, dando ao
distribution_distance_score uma perna de referência trace-backed real.
A métrica de calibração nomeada reporta a distância 1-Wasserstein (movedor-de-terra)
por-característica entre as distribuições de características de caracterização de um perfil
calibrated-synthetic e as de suas instâncias de referência trace-backed. As distâncias
por-característica são agregadas a uma única pontuação de distribution-distance; uma pontuação acima
do limiar de divergência-máxima (padrão 0.25) significa que o perfil derivou longe demais
de sua referência e falha na calibração. A métrica é reproduzível a partir das instâncias
materializadas e do traço de referência nomeado.
from dispatchatlas.bench import (
build_smoke_catalog,
metrics_from_instance_set,
distribution_distance_score,
)
catalogs = {c.config.profile_id: c for c in build_smoke_catalog()}
synthetic = metrics_from_instance_set(catalogs["cloud-edge-capacity"])
reference = metrics_from_instance_set(catalogs["workflow-dag"])
score = distribution_distance_score(
synthetic, reference, reference_trace_id="google-cluster-data"
)Estratificação e seleção de subconjunto
difficulty_score agrega os descritores de contenção, sobrecarga,
profundidade-de-dependência, e sensibilidade-do-solver em uma pontuação de dificuldade
normalizada, e stratify_instances agrupa as instâncias materializadas em estratos de
dificuldade baixa, média, e alta. select_benchmark_subset escolhe um subconjunto
determinístico filtrado por família, classe de perfil, e estrato de dificuldade, ordenado
por identificador de problema de modo que a seleção é reproduzível.
Catálogo de fumaça
from dispatchatlas.bench import build_smoke_catalog, smoke_benchmark_provider
catalogs = build_smoke_catalog(root_seed=20260527)
provider = smoke_benchmark_provider(root_seed=20260527)
first_problem = provider.get_problem(provider.list_problem_ids()[0])O catálogo de fumaça empacotado inclui as duas famílias de distributed-computing (cloud/edge de tarefa-independente e workflow DAG) junto às quatorze famílias genéricas de escalonamento (escalonamento de máquinas, job-shop, job-shop flexível, flow-shop de permutação, flow-shop de setup dependente-de-sequência, RCPSP, open-shop, flow-shop híbrido, flow-shop de permutação distribuído, flow-shop sem-espera, flow-shop com-bloqueio, flow-shop de montagem distribuído, flow-shop de permutação multi-objetivo, e RCPSP/max) como pares co-iguais. Cada instância é determinística a partir da semente raiz e usa metadados de gerador respaldados-por-citações enquanto permanece rotulada como material de desenvolvimento de fumaça.
Catálogos completos candidatos
Os catálogos de campanha-completa candidatos usam o mesmo caminho de materialização com contagens de problema configuradas maiores e rótulos de evidência mais estritos:
from dispatchatlas.bench import build_full_catalog, full_benchmark_provider
catalogs = build_full_catalog(root_seed=2026052713, problem_count_per_profile=30)
provider = full_benchmark_provider(
root_seed=2026052713,
problem_count_per_profile=30,
)Esses catálogos de grau de evidência de campanha-completa são respaldados-por-citações, caracterizados, ligados-por-hash, e ainda marcados como evidência de calibração até que os gates de campanha, divulgação, e publicação promovam afirmações específicas.